Anthropic has published the methodologies behind three metrics it already tracks internally: AI-led research and development, oversight of autonomous AI agents, and compute allocation. The disclosure follows a call by chief executive Dario Amodei for leading model makers to coordinate on slowing the pace of frontier development, and it is meant to let other frontier labs measure the same things. The company says the gap between what frontier labs know and what the public knows should be as small as possible.
What the three metrics measure
The first metric is an index of how much of Anthropic's own R&D work is done by Claude. The company reports that the model is not operating fully autonomously in any subset of that work. The second metric covers oversight of AI agents: Anthropic has built a system that can supervise agents and intervene in the actions they take inside its internal systems. That system shows the company is currently running about 30,000 autonomous agents in its computing environments, most of them involved in R&D. The third metric is a regular snapshot of compute allocation, the resource Anthropic calls the primary fuel AI runs on.
The first such snapshot covers July 13 to July 20 and shows that around 6% of total compute capacity went to AI safety, with a further 12% dedicated to AI-driven R&D focused on safety. Anthropic argues that compute is among the most verifiable inputs to the AI R&D process, which makes it a possible lever in any future pacing effort. The company says the combined measurements give third-party evaluators a starting point for judging how fast AI is really accelerating and what it is capable of, and that it intends to keep releasing them.
The disclosure fleshes out a three-step plan Amodei published at the weekend: measure the development of AI, report on it publicly, and give society an opportunity to decide how to use that information. Amodei framed the goal as curtailing the rapid pace of frontier model development without sacrificing the commercial advantage of the United States' lead in AI. The plan has drawn backing from peers including OpenAI chief executive Sam Altman, SpaceX chief executive Elon Musk and Google DeepMind chair Demis Hassabis.
What this means for business
For companies that buy or build on frontier models, the practical change is the arrival of comparable numbers. Until now, assessments of how quickly AI capabilities were moving came mostly from vendor statements and benchmark results. An index of AI-led R&D, a count of running agents and a breakdown of compute allocation give procurement and risk teams something to ask vendors about directly. A small company evaluating a single model provider can use the agent count as a rough signal of how much autonomy is already in production use; a large enterprise running its own agent fleets can compare its oversight practices against the intervention system Anthropic describes.
The limits matter as much as the figures. These are Anthropic's own measurements of its own systems, not an industry standard, and the company has not said how the R&D index is calculated or what a high or low value would mean. The 6% and 12% compute figures come from a single week in July and cover one lab. Nothing here shows that other frontier labs will adopt the same definitions, and nothing obliges them to publish. A business weighing an AI vendor should ask which of the three metrics the vendor tracks, over what period, and whether the numbers are audited or self-reported.
The marker to watch is whether other frontier labs publish the same three metrics with the same definitions. If they do, buyers get a common yardstick for comparing how fast models are advancing and how much oversight sits behind them; if the disclosure stays with one company, the figures remain a single vendor's account of its own progress.
