AERIOXFLUX
Frontier Labs
Frontier Labs · benchmarks safety

Claude Now Leads 26% of Anthropic's Own Research

In February it was under 1%. Anthropic published the methodology, the automation scale, and the number nobody else reports — plus a 59% agreement rate that undercuts all of it.

Flux Desk·2026-09-19·5 min read

The Anthropic Institute published a prototype R&D Automation Index on September 17. Its headline finding: as of August 2026, Claude "leads" 26% of Anthropic's AI research and development work, up from under 1% in February.

More than 90% of that work sits at or above the "collaborates" level. The fraction operating with no human in the loop is zero.

That last number is the one people will skip past, and it is the most important one in the report.

The scale is the contribution

Anthropic scored tasks on an AL0–AL5 automation scale developed by Epoch AI. AL0 is no AI involvement. AL3 is Claude completing large portions under close direction. AL4 — "leads" — is Claude handling most of a task end-to-end from a high-level prompt with a human supervising. AL5 is fully autonomous operation with no human oversight.

The 26% figure is the AL4-and-above share. AL5 was measured at 0% across every subset.

The methodology is unusually legible for a self-reported metric. For each week of July 2026, a Claude research agent reviewed a random 20% sample of staff from every department in the model R&D loop and enumerated what they worked on, producing roughly 15,000 granular tasks. Claude then organized those into a hierarchical tree of 542 nodes, 378 of them leaves — things like "eval platform defect diagnosis and fixes." A separate Claude judge assigned automation levels. Employees self-rated without seeing the model's evidence.

Two companion measurements shipped alongside it. On oversight: roughly 30,000 agents running research and engineering work simultaneously, online monitors covering 100% of agent actions before execution, and a blocking rate of 0.002% — about 1 in 47,000 decisions out of more than a billion. Offline monitors flag 1–2 transcripts per 1,000 actions, and roughly 50 transcripts a week escalate to human review.

On compute, from a single week sampled July 13–20: 6% of AI R&D compute and 12% of AI-driven R&D compute went to safety work. Anthropic calls these conservative.

The 59% is the caveat that matters

Anthropic states plainly that no outside party has checked the figures. The internal validation numbers are more interesting than that disclaimer.

The model's automation ratings matched employees' self-ratings exactly 59% of the time. Employees matched each other exactly only 35% of the time. Model and human ratings landed within one level of each other 97% of the time.

Read those three together. The humans disagree with each other more than the model disagrees with the humans, which is either evidence the model is a reasonable judge or evidence that the underlying construct — "how automated is this task" — is too fuzzy to pin to a discrete level. Both readings survive the data.

Anthropic names the deeper problem itself: the judge model "could make the same kinds of errors as the model it is checking." A Claude agent enumerating tasks, a Claude judge scoring them, and a report about how much Claude does is a measurement loop with one vendor inside every stage.

The remaining constraints are worth stating. The index covers only Anthropic's most-used internal platform. The compute figure is a one-week sample the company itself says is too short for a trend. Mixed-purpose work was classified as capabilities rather than safety, and safeguard classifiers were excluded from the compute math. Compute measures spending, not effort.

Why publish this at all

Because "is recursive self-improvement happening" has been argued entirely on vibes for three years, and 26% AL4 with 0% AL5 is a falsifiable claim.

It is also a calculated disclosure. Anthropic has been arguing publicly — in Amodei's September 12 pacing essay and in the Accenture evaluator partnership signed six days later — that frontier development needs external measurement. Publishing an unflattering-looking number first is how you establish that the measurement should be yours.

And the 6%-to-safety figure is the one that will be quoted against the company for years. Anthropic published it anyway, which is either commendable or a controlled release of the criticism it expects regardless.

The number that isn't in the report

Twenty-six percent of Anthropic's R&D. Not the industry's.

There is no comparable figure from OpenAI, Google DeepMind, xAI or Meta, because none of them publish one. The AL0–AL5 scale only becomes useful if it becomes shared. Right now Anthropic has unilaterally handed regulators and critics a metric that exactly one company reports — which means every future comparison, fair or not, uses Anthropic as the reference case.

That is a real cost, and it is the strongest evidence that the publication was sincere rather than tactical.

What to watch

Whether any other lab adopts AL0–AL5. If two labs report on the same scale, it becomes a standard. If none do, it stays a marketing asset.

The next compute snapshot. One week is not a trend. Four quarters of the safety-compute share is the number that tells you whether 6% was a floor or a ceiling.

Whether the 0% at AL5 holds. That is the line the entire report is built to defend, and it is the one that will be tested first.

Third-party verification. Anthropic says outside evaluators are planned but not yet embedded in this measurement. Until one is, this is a company grading its own homework and showing the work.

#anthropic#automation#recursive-self-improvement#measurement#claude

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.