Anthropic is paying $1 billion to have someone watch it train models

Accenture’s Faculty unit gets access comparable to an employee’s inside Anthropic, a step toward the slowdown Anthropic’s own chief executive proposed weeks earlier.

Abstract EMRGNG cover image for a story about Anthropic, Accenture

Anthropic named Accenture’s Faculty division as its first embedded AI evaluator on 18 September, with each company expecting to invest at least $1 billion over five years, a combined $2 billion commitment that Anthropic is funding directly. The arrangement puts independent evaluators to work inside Anthropic itself, rather than testing finished models from the outside after training is done.

Faculty’s evaluators will have access comparable to an employee’s, letting them observe training runs, follow deployment decisions and talk directly with staff, while running red-teaming, alignment assessments, safeguard testing and incident reporting on an ongoing basis. Anthropic frames the structure as one step toward a commitment made in chief executive Dario Amodei’s essay calling for the industry to pace its own development, published in early September.

The deal is explicitly non-exclusive. Anthropic says it is in discussions with the research nonprofit METR and other third parties about similar embedded arrangements, suggesting the company wants a standing roster of internal watchdogs rather than a single partner. It is a notable structural choice for a company competing directly with OpenAI and Google on model capability to also be the one funding the people paid to scrutinise it.

Anthropic says embedded evaluators do not reduce its own accountability, only make it more verifiable, and that responsibility for model safety still sits with the company. What is not addressed is what happens when an evaluator Anthropic is paying flags something Anthropic disagrees with, whether those findings ever become public, or who has the final say when the funder and the watchdog disagree.

Read more here.

More from EMRGNG