External safety teams could gain deep access to internal systems and the ability to report incidents publicly, if the companies follow through on their commitments. Anthropic’s CEO Dario Amodei proposed embedding independent researchers inside the firm, and OpenAI’s CEO Sam Altman signalled a similar approach. Both companies described giving outside groups access to internal models and the power to assess alignment and publish findings, and named evaluators such as METR and Redwood Research as examples that could receive unprecedented access.

Researchers welcomed the move in principle, but emphasised that raw access will not automatically produce independence. Alexander Meinke, head of research at Apollo Research, said AI developers should be able to state whether a model ever tried to undermine its alignment training, adding, “the answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we've seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.”

Evaluators want more than tests of finished products. Independent teams say intermediate training snapshots and logs, often called checkpoints, reveal when concerning behaviour first appeared. FAR.AI chief Adam Gleave said those records would allow outsiders to compare versions, inspect reward systems applied after training, and verify evaluation transcripts and run logs companies cite when describing performance.

Key operational details remain unsaid. Neither Anthropic nor OpenAI has disclosed which evaluators would be invited, how many would be embedded, when embedding would begin, what internal systems each team could review, or what publication rights teams would have. Amodei’s outline does include a formal right for embedded teams to “publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive, without editorial control by Anthropic,” but comparable specifics have not been published for either firm.

The absence of hard rules matters because previous outside reviews were constrained by time, scope and confidentiality. FAR.AI has declined contracts when developers demanded excessive editorial control. When METR and Redwood investigated the Hugging Face incident, they had roughly a week on premises and later said those limits helped prevent confident conclusions. Apollo Research reports it was given three days to test GPT-6 Astra, a window its team said was too brief to demonstrate alignment definitively.

The proposal represents a clear shift from firms that historically limited external review to pre-release tests. For the initiative to be meaningful, companies must name partners, define the scope and timing of embedded work, and set enforceable disclosure rules. Evaluators say only then will the industry’s promise of inside oversight be testable in practice.