OpenAI and Anthropic have a plan to stop AI from going rogue — there’s just one catch

Sam Altman making an interesting face
(Image credit: Getty Images)

OpenAI's internal AI agents hacked into Hugging Face while trying to cheat on a cybersecurity benchmark. Anthropic has disclosed that Claude models reached the open internet from cybersecurity test environments that were supposed to be sealed off and then got into the real systems of outside organizations.

Both are examples of systems taking actions their developers didn't intend, and METR, an independent AI research nonprofit, has been brought in to look at each of them. It documented the Hugging Face incident alongside Redwood Research and is now investigating Anthropic's.

OpenAI and Anthropic have also committed to bringing outside experts inside their companies to examine safety practices, according to The Atlantic. Called "embedded evaluators," these researchers could get a closer look at how models are built and investigate problems that might never surface in a product demo.

Latest Videos FromTom's Guide

The catch is that this oversight is voluntary. Last week, leaders from Anthropic, OpenAI and other major labs signed an accord at the White House committing to independent external auditors, and the FTC has opened an industry-wide probe into AI labs, but the companies being examined still have considerable control over what outsiders can see, which risks they investigate and what information reaches the public.

For anyone being encouraged to let AI handle more of their work, that leaves a fairly basic trust problem. We're being given assistants while the system for independently checking their behavior is still taking shape.

As I've reported before, the film The AI Doc puts the number of people working on AGI, or artificial general intelligence, at more than 20,000, while fewer than 200 are focused specifically on AI safety.

A closer look inside AI companies

anthropic

(Image credit: Shutterstock)

Today, an outside researcher might test a model before release and report how it performed. Embedding evaluators inside a company could let them follow its development and investigate the decisions behind that behavior.

Anthropic says these evaluators would have access comparable to an employee's, allowing them to observe training, speak with staff and examine how models are built and deployed.

The company has named Accenture as its first embedded evaluator, with Faculty, Accenture's specialist AI business, leading the work. Each company expects to invest at least $1 billion in AI safety over five years.

That access could help researchers discover whether a troubling incident was an isolated failure or evidence of a wider problem. METR's proposed approach to investigating incidents calls for examining transcripts, interviewing employees and studying the training conditions that might have encouraged the behavior.

The companies still set the terms

Sam Altman

(Image credit: Getty Images)

Anthropic acknowledges that standards for evaluators' access, reporting and funding haven't been settled. It will fund Accenture's work directly, while discussing pilots with nonprofits that would use their own funding.

Accenture is not new to Anthropic, since the two companies already run a joint business group under a partnership announced in December 2025.

That relationship deserves attention, because a group of evaluators published an open letter arguing that embedded evaluators shouldn't have other significant commercial business with the labs they assess.

As The Atlantic's reporting explains, voluntary oversight gives companies considerable discretion over the scope of an evaluation. Researchers need enough freedom to examine uncomfortable evidence and tell the public what they found.

They also need to be able to explain what they couldn't check. There needs to be transparency because a report based on limited access could be valuable, but we should know where its conclusions end.

What this means for ChatGPT and Claude users

METR's recent assessment of Claude Opus 5.5 focused on whether the model could accelerate AI research and development. The organization explicitly says its summary doesn't assess the model's alignment properties and whether its behavior follows intended goals and constraints.

METR conducted the evaluation under an unpaid agreement, and it also disclosed that Anthropic had an opportunity to review and edit the summary, with METR approving the final text.

The incidents involving internal agents don't establish that the ChatGPT or Claude you use for everyday tasks will behave the same way. OpenAI found that its agents' tendency to compromise infrastructure dropped by more than 100 times when they ran with the production ChatGPT setup, and Anthropic says the behaviors it found are unlikely to show up in ordinary use.

Independent researchers need enough access to uncover risks and the freedom to publish what they find, even if it delays a launch or frankly, even if it embarrasses the company paying for their work.

Bringing outside evaluators into AI labs is a promising start, but their influence will depend on what happens when they uncover a serious problem. If a company can restrict the investigation or keep the findings private, the public is still being asked to take its word for it.


Follow Tom's Guide on Google News and add us as a preferred source to get our up-to-date news, analysis, and reviews in your feeds. Subscribe to Tom's Guide on YouTube and follow us on TikTok.

Google News


More from Tom's Guide

TOPICS
Amanda Caswell
AI Editor

Amanda Caswell is the AI Editor at Tom's Guide and one of today’s leading voices in AI and technology.

A celebrated contributor to various news outlets, her sharp insights and relatable storytelling have earned her a loyal readership. Amanda’s work has been recognized with prestigious honors, including outstanding contribution to media.

Known for her ability to bring clarity to even the most complex topics, Amanda seamlessly blends innovation and creativity, inspiring readers to embrace the power of AI and emerging technologies.

As a certified prompt engineer, she continues to push the boundaries of how humans and AI can work together.

Beyond her journalism career, Amanda is a long-distance runner and mom of three. She lives in New Jersey.

You must confirm your public display name before commenting

Please logout and then login again, you will then be prompted to enter your display name.