OpenAI chief executive Sam Altman committed on September 12, 2026 to giving independent evaluators employee-like access to the company's systems, matching a pledge Anthropic CEO Dario Amodei announced earlier the same day. In a post on X, Altman wrote that the idea is a good one and that OpenAI will do the same, adding that more details will follow soon. The two largest US frontier labs now publicly accept outside reviewers inside their own offices, which changes how safety claims in this market can be checked.
What the two labs committed to
Amodei's pledge is the first step of a three-step plan set out in an essay titled We Must Pace the Frontier, dated September 2026. Anthropic commits unilaterally to give third-party evaluators permanent, employee-level access to its systems so they can verify adherence to safety measures, report incidents, and assess how models behave during training, not only after release. The essay names METR as an example of such a reviewer. Amodei wrote that Anthropic intends to invite an embedded external review team with desks in its offices, access badges and company laptops, plus access to workspaces, tools and permissions mostly comparable to what internal risk-assessment teams hold. Exceptions are reserved for cases where the law or contracts require it, or to protect private information of customers and partners.
The mechanics matter more than the wording. Reviewers would hold the right to publish key findings about risk levels, incidents, practices and the access they received, without editorial control by Anthropic. The company keeps a narrow ability to redact security-sensitive, legally privileged, commercially sensitive or third-party confidential material, but cannot redact findings merely because they are unfavorable, and reviewers may state publicly if a redaction removed something important to their conclusions. Amodei described the arrangement as going far beyond the practices of any AI company and urged other frontier labs to follow. Steps two and three of the essay call for coordination on common safety standards among frontier companies in democratic countries, and then global coordination involving authoritarian governments.
OpenAI's own record explains why the pledge is not a sudden turn. In an August 18, 2026 post, the company said it had temporarily slowed the pace of scaling, including a two-week pause in reinforcement learning training on its latest models intended for deployment, while it hardened and red-teamed research environments and expanded monitoring coverage. Its largest planned frontier reinforcement learning run remained on hold while smaller-scale training and evaluations continued. The same post cited the OpenAI-Hugging Face incident and preliminary evidence that the upcoming Astra model may meet the Critical cybersecurity capability threshold under its Preparedness Framework, with monitoring overhead estimated at roughly 20 percent of the inference compute being monitored. OpenAI already runs a third-party assessment program described on November 19, 2025: independent evaluations of frontier capability and risk, methodology reviews, and subject-matter expert probing, with assessors under non-disclosure agreements and publications reviewed by OpenAI for confidentiality and factual accuracy.
What this means for business
For companies that buy or deploy AI, the practical change is the arrival of a documented outside check on safety claims. Until now, a procurement team comparing vendors had little more than self-reported evaluations and voluntary model cards. An embedded reviewer with badge-level access produces findings that can be cited in a vendor review, an internal risk register or a board paper, and the right to publish without editorial control makes those findings harder to dismiss as marketing. The effect is uneven: a large enterprise with a formal AI governance function can fold such reports into existing audit cycles, while a small company will mostly benefit indirectly, through clearer public statements about incidents and risk levels that it could never obtain on its own.
The limits are equally concrete. Anthropic's commitment is unilateral and covers one company; Altman's reply promises more detail later and does not yet specify scope, timing or which evaluators OpenAI would admit. Redaction rights remain, so a published finding may be incomplete, and reviewers can flag that only if they choose to. OpenAI's existing assessors sign non-disclosure agreements and their publications pass an internal review for confidentiality and factual accuracy, a different model from the one Amodei describes. Buyers should therefore ask vendors which outside party has ongoing access, what that party may publish, who reviews it before release, and whether compensation depends on results. The news does not mean frontier models are now independently certified, nor that safety claims have been verified by a regulator.
The marker to watch is the first published report from an embedded team. Amodei wrote that Anthropic intends to invite its review team in the near future, and Altman said OpenAI will share more soon; if a named evaluator such as METR or the newly announced Open Alignment Initiative, led by Hugging Face co-founder Thomas Wolf, publishes findings that a lab cannot edit, the model becomes verifiable rather than declaratory. If instead the access arrives without public output, the pledge stays a statement of intent, and procurement teams should keep treating safety claims as unverified input to their own testing.
