Outside evaluators will receive employee-level access inside Anthropic under a commitment made in Dario Amodei's essay We Must Pace the Frontier. The first partner is Accenture, with Anthropic and Accenture each planning to invest at least $1 billion over five years to build evaluation capacity. The arrangement moves independent oversight of frontier models from a proposal to an operating practice that business users of AI will need to track.

Anthropic gives Accenture-led evaluators inside access to its models

Accenture partnership and METR talks

The work will be led by Faculty, Accenture's specialist AI business. The remit includes evaluating and red-teaming models, conducting alignment assessments and testing safeguards. Anthropic cites Accenture's experience deploying AI for businesses and governments across industries as a practical view of enterprise use. Anthropic will fund Accenture's work directly. Separately, it is in dialogue with METR and other nonprofit evaluators to pilot embedded evaluation with their own funding. The deal is non-exclusive, with additional evaluators expected in the coming weeks.

Embedded evaluation means outside specialists work inside the lab with access equivalent to highly privileged employees, with exceptions for protecting data, while keeping an outside perspective and reporting on what they find. Minimum standards outlined in a public letter led by Geoffrey Hinton, Stuart Russell and Arvind Narayanan call for meaningful independence of such evaluators. Payments cannot depend on findings, there should be no ownership, governing or commercial relation and no editorial control. Conflicts must be disclosed, differing viewpoints included, and evaluators shielded from retaliation.

OpenAI has followed Anthropic in committing to such evaluators and added a call for international coordination. The central dispute is who selects and pays for the work, since experienced people usually built their expertise inside the top labs. Established audit models still struggle with the revolving door and the source of funding. Gabriel Weil argued in July that developers should not hire their own referees and proposed mandatory liability insurance. Discussion in the source covers $100 billion or $1 trillion in coverage, publication of its cost and disclosure of pricing risk factors, alongside mention of a $2.2 trillion valuation and recent AI hacking incidents.

What this means for enterprise AI buyers

For companies that buy or build on frontier models, a standing evaluation layer could make safety claims easier to compare across vendors. An enterprise-oriented evaluator brings knowledge of how models are actually deployed in regulated workflows, procurement reviews and operational controls. A mix of corporate and nonprofit evaluators would widen the range of findings, from usability of safeguards to alignment risks. Large organizations may use such reports in vendor due diligence, while smaller firms are more likely to rely on summarized results and procurement checklists.

The setup leaves open questions that matter for trust in any report. Anthropic previously partnered with Accenture and will fund this work directly, which creates a conflict of interest even with formal independence rules. Multiple sources of oversight can help, but a lab could also highlight the most favorable review and discount sharper criticism. Observers cited in the source divide on this point, with skepticism about consulting expertise on one side and Faculty's prior evaluation work on the other, plus a request for Accenture to publish an example of a frontier-model assessment.

The marker to watch is whether Anthropic adds METR and other nonprofits in the coming weeks and publishes methods and findings from both tracks. Parallel pilots with independent funding, plus Accenture engagements with other developers, would show an ecosystem rather than a single bilateral deal. Absence of those steps would leave the model dependent on one paid relationship.