Anthropic reported that its Claude models lead 26% of the company's AI research and development work as of August 2026, up from under 1% in February. The figure comes from a prototype R&D Automation Index published in an Anthropic Institute post on September 17, 2026, the first attempt by a frontier lab to put a number on how much of its own AI work is done by AI. For businesses watching where model development is heading, the index offers a rare look at the internal mechanics of a lab rather than another benchmark score.
What the index measures
The index maps the full range of AI R&D work done at Anthropic, scores how automated each task is, and combines the scores into an aggregate measure. The scoring uses an Automation Level scale developed by Epoch AI, running from AL0, meaning no AI involvement, to AL5, where AI operates fully autonomously with no human in the loop. At AL3, AI collaborates, performing large portions of a task under close human direction; at AL4, AI leads, carrying most of a task from a high-level prompt to completion while a human supervises. Anthropic reported that the share of work at or above the collaborates level is above 90%, and that Claude is not operating fully autonomously for any measured subset of AI R&D work.
The task catalogue was assembled from the bottom up using work records such as Slack and internal documentation. For each week of July 2026, a Claude research agent reviewed each randomly sampled person's week, covering 20% of staff from every department that makes up the model R&D loop, and listed the tasks they worked on, yielding a flat list of roughly 15,000 granular tasks. Claude then organized those tasks into a hierarchical tree of 542 nodes, 378 of them leaves such as eval platform defect diagnosis and fixes. That tree is frozen so every measurement runs against the same basket of work.
For each node, a Claude agent researches how that kind of work is done across the company, and an independent Claude judge assigns one of six automation levels, restricted to evidence from the month being rated or earlier. Tasks are weighted by person-time, so categories where more staff effort goes carry more weight. Anthropic checked the judge's ratings against staff who own the relevant work areas, who rated without seeing the models' evidence or judgments, and reported model-versus-human exact agreement of 59% versus 35% human-versus-human, with ratings within one level of each other 97% of the time. Stated limitations include the frozen basket, which captures automation of existing work without registering new kinds of work; a comparison of tasks arriving from February through July 2026 against a January 2026 basket found no rise in novel tasks, and Anthropic plans to rebuild the basket periodically and re-version the published numbers.
What this means for business
For companies adopting AI agents, the index is a template for measuring automation without relying on vendor claims. The method weighs tasks by person-time, which means it tracks where staff effort actually goes rather than counting tools deployed. A small company can run a simplified version of the same exercise: list recurring tasks, score each on the AL0 to AL5 scale, and check the scores against the people who own the work. The 59% exact agreement between model and human raters suggests such self-assessment is directionally useful but not precise enough to base headcount decisions on a single quarter.
The oversight figures published alongside the index matter for anyone running agents in production. As of August 2026, approximately 30,000 agents were doing research and engineering work at any one time on Anthropic's most-used internal platform. Online monitors check 100% of these agents' actions before execution, usually within seconds; of more than a billion agent decisions analyzed over August 2026, 0.002%, about 1 in 47,000, were blocked, and humans review any blocked actions within one week. Offline monitors ingest 100% of actions after the fact and flag roughly 100,000 transcripts per week, with approximately 50 highest-priority flags per week escalated to human review. Buyers of agent platforms should ask vendors which of these two layers they operate, what share of actions is checked before execution, and how blocked actions are resolved.
Anthropic also measured compute allocation from July 13 to July 20, 2026, sorting every workload into categories: about 6% of compute going to AI R&D was allocated to safety, and about 12% of compute going to AI-driven AI R&D went to safety. The company describes both estimates as deliberately conservative, counting tokens that advanced capabilities as much as safety as AI R&D and excluding safeguards classifiers. A prompted Claude classifier sorted the week's almost 10,000 research training and evaluation runs using a roughly 14% sample weighted toward the largest compute users, agreeing with human reviewers within one or two percentage points. The stated limitations are that a single week demonstrates feasibility without establishing a trend, that workload labels are best-effort and unverified, and that compute share measures spending rather than the amount of safety work performed. The marker to watch is whether Anthropic rebuilds the basket and publishes a second index; if the 26% leads share keeps climbing while the safety compute share stays near 6%, the gap between capability automation and oversight investment becomes the number that matters for regulators and enterprise buyers alike.
