Anthropic said on September 18, 2026 that it is partnering with Accenture on independent evaluation of frontier AI, with each company expecting to invest at least $1 billion in building capacity in the area over the next five years. The work will be led by Faculty, Accenture's specialist AI business, and covers evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards. For companies that buy AI systems, the arrangement is a test of whether safety claims can be checked by an outside party rather than taken on trust.

Anthropic Taps Accenture's Faculty for Embedded AI Model Evaluation

Who runs the evaluation and on what money

The partnership is non-exclusive on both sides. Anthropic said it will work with other evaluators it plans to announce in the coming weeks, expects frontier labs to work with several organizations at once, and accepts that Accenture will work with other AI developers in similar capacities. Funding is arranged per evaluator: Anthropic will pay Accenture directly, while talks with METR and other nonprofit evaluators cover pilots of embedded evaluation elements using those organizations' own funding. Anthropic said there are no standards yet for what information embedded evaluators should receive or how they should report findings, and no settled system for funding independent evaluation. Long term, the company argues funding should come from pooled or government sources, as set out in its Advanced AI Framework in June; since neither exists, different evaluators will work under different arrangements.

The mechanics differ from the usual external audit. Where existing external evaluators work outside the companies they assess, embedded evaluators sit inside the AI company with access comparable to an employee's. That position lets them watch models take shape during training, follow the decisions that govern how models are built and deployed, and speak directly to employees. From there they can assess how the company operates, verify that it is keeping its safety commitments, and identify blind spots, as well as report incidents and give the public a more informed account of benefits and risks. Anthropic stated that independent embedded evaluators leave its accountability unchanged while making that accountability more verifiable, and that the safety of its models remains its own responsibility.

The arrangement follows a September 2026 essay by CEO Dario Amodei, titled "We Must Pace the Frontier", in which he proposed a three-step plan: embedded evaluators, coordination among frontier AI companies in democratic countries, and global coordination. Anthropic committed unilaterally to the first step and called on governments to require other frontier companies to match it. Amodei wrote that the pace at which AI models improve must be slowed, that progress will still seem fast, and that the time gained should be used wisely. The essay describes the intended setup: an external review team with desks in Anthropic's offices, access badges, company laptops, and access to workspaces, tools, and permissions mostly comparable to internal risk-assessment teams, with exceptions where law or contracts require it or to protect customers' and partners' private information. Reviewers would hold the right to publish key findings about risk levels, incidents, practices, and the access they received or did not receive, without editorial control by Anthropic; the company keeps a narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but cannot redact findings merely because they are unfavorable, and reviewers may say publicly if a redaction removed something important to their conclusions. The essay points to banking, where regulatory supervisors are sometimes embedded alongside employees.

What this means for companies deploying AI

For buyers of AI systems, the practical change is the appearance of a second source of information about model risk. Until now, an enterprise assessing a vendor relied on the vendor's own safety documentation, its model cards, and its incident reports. An embedded evaluator with employee-level access can confirm or contradict those statements from inside, and the right to publish findings without editorial control means the result can reach the buyer even when it is unfavorable to the vendor. In procurement this matters most in regulated sectors, where a risk committee has to justify a decision on paper: an external assessment of safeguards and alignment becomes an argument in that file. The effect is uneven. A large bank or insurer with a formal AI review process can fold such reports into existing vendor due diligence, while a small company without a risk function will mostly notice the change indirectly, through contract terms and the questions vendors start asking about how their models are used.

What the news does not settle is how comparable these assessments will be. Anthropic itself says there are no standards for evaluator access or reporting, and no funding model, so two evaluators may work under different rules and produce findings that cannot be lined up against each other. The contract described in the essay gives Anthropic a narrow right to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, and reviewers may state publicly when a redaction removed something important to their conclusions; that disclosure is the main check on the carve-out, and it depends on the reviewer using it. The partnership is also non-exclusive, so an assessment by Faculty says nothing about models from other developers, and Anthropic continues to train and release frontier models while the evaluation framework is being built. A company choosing a vendor should therefore ask what exactly the evaluator could see, what it was not allowed to see, whether it could publish without approval, and who paid for the work.

The marker to watch is the set of additional evaluators Anthropic says it will announce in the coming weeks, together with the first published findings from the Accenture work. If those reports name specific risk levels, incidents, and access limitations, and if other frontier labs sign comparable arrangements, embedded evaluation becomes a normal input into enterprise procurement rather than a single vendor's experiment. If the announcements stay at the level of partnerships and the findings do not appear, buyers should keep treating safety claims as unverified.