Anthropic Hired a Consulting Firm to Watch Itself
The first embedded evaluator under Dario Amodei's pacing proposal is Accenture, not METR — and the two companies will spend roughly $2 billion over five years defining what independent oversight means.
On September 18, Anthropic and Accenture announced that Accenture evaluators will be embedded inside Anthropic with access "comparable to an employee's" — observing model development, following deployment decisions, and interacting with staff directly. The work will be led by Faculty, Accenture's specialist AI business. Each company expects to invest at least $1 billion over five years, roughly $2 billion combined.
The scope is red-teaming, alignment assessments, and testing model safeguards. The arrangement is non-exclusive on both sides: Anthropic says more evaluators will be named in the coming weeks, and Accenture is free to work with other developers. Separately, Anthropic says it is in dialogue with METR and other nonprofit evaluators about piloting embedded evaluation funded by the nonprofits themselves.
Accenture stock rose roughly 8% after hours.
Six days from essay to signed contract
Dario Amodei published "We Must Pace the Frontier" on September 12. Its first commitment was embedding third-party evaluators inside labs. The first one was signed six days later.
That is fast even by this industry's standards, and the speed is the most informative thing about the announcement. A proposal that goes from essay to billion-dollar partnership inside a week was not a proposal. It was a pre-announcement.
This matters because the sequencing determines who writes the rules. There is currently no standard for what an embedded evaluator gets to see, what it may publish, whether it can halt a deployment, or who pays. Anthropic has now answered all four questions unilaterally: employee-level access, unspecified disclosure rights, no stated veto, and the lab pays half.
Whoever sets the first template usually sets the norm. Anthropic knows this.
Why Accenture is the argument
The critique writes itself, and TechCrunch reported that the choice "surprised many AI watchers." Accenture is a consulting firm with a market capitalization in the tens of billions whose business is selling AI implementation services to enterprises. METR and Redwood Research are safety-native organizations whose entire reason for existing is evaluating frontier models adversarially.
Anthropic picked the consultancy and is paying it.
The structural objection is not that Accenture is incompetent — Faculty has real capability and Accenture can field evaluators at a scale no nonprofit can match. The objection is that an evaluator selected by the lab, paid by the lab, and commercially dependent on the broader AI buildout is not independent in the sense the word usually carries. An auditor whose client can fire it and whose sector-wide revenue depends on AI adoption continuing has a conflict that no access agreement resolves.
Anthropic's implicit answer is capacity. There are perhaps a few dozen people on Earth qualified to do adversarial frontier evaluation full-time, and they do not all work at METR. If embedded evaluation is going to exist at the scale of a lab running 30,000 concurrent agents, somebody has to hire and train hundreds of people, and that is a thing consulting firms are genuinely good at.
Both arguments are true. Which one dominates depends on a detail neither company disclosed: what Accenture is permitted to publish, and whether Anthropic can stop it.
The nonprofit track is the tell
Note the structure. Accenture gets a commercial contract with billion-dollar investment on both sides. METR and other nonprofits get a "dialogue" about piloting embedded evaluation using the nonprofits' own funding.
That is two tiers. The well-resourced evaluator is paid; the independent evaluator pays its own way. If you believe the conflict-of-interest critique, the tiering makes it worse — the organizations least compromised by the arrangement are the ones being asked to self-fund access.
If you believe the capacity argument, the tiering is simply what it looks like when a lab buys services from a vendor and separately offers observation to researchers. Neither reading is unreasonable. Both should make you want the disclosure terms.
The state caught up in two days
On September 18 — the same day — California Governor Gavin Newsom signed an executive order directing state agencies toward requiring frontier AI companies to embed independent verification organizations onsite for regular audits.
So within a single week, embedded evaluation went from a lab CEO's essay, to a private commercial partnership, to a directive to write it into state policy. The industry proposed the mechanism and a regulator adopted the vocabulary almost immediately.
That is the actual competitive dynamic here, and it explains the hurry. A voluntary framework you design is much better than a mandatory one somebody else designs, and the window to demonstrate that voluntary works closes the moment a state statute lands.
What to watch
The disclosure terms. Everything turns on whether Accenture can publish findings Anthropic dislikes. Until that is public, this is an unaudited claim about auditing.
Whether METR takes the deal. A nonprofit accepting self-funded embedded access legitimizes the two-tier structure. One declining publicly would be the loudest available critique.
Who OpenAI and Google DeepMind hire. If they pick consultancies too, embedded evaluation becomes a professional-services category within a year, with the accompanying accreditation, billing rates, and race to the bottom.
Whether California's order names acceptable evaluators. A state registry of approved auditors — which AB 1405 already contemplates — would take the selection decision away from the labs entirely. That is the version of this that has teeth.
