Anthropic & Claude

Five conditions for embedded evaluators, signed by Hinton and Russell

3 min read AI-generated

Evaluators should get the same access as highly privileged employees — plus protection from lawsuits filed by the company they are auditing.

Featured image for "Five conditions for embedded evaluators, signed by Hinton and Russell"

One day after Anthropic named Accenture as its first embedded evaluator, more than a hundred researchers, evaluators and industry figures published an open letter. It welcomes the idea, then spells out how you can tell whether anyone means it.

The five minimum conditions

The AI Evaluator Forum letter lists five things without which embedded evaluation cannot be credible.

One, independence. Evaluation organisations must not be owned by the lab. They should have no significant other commercial business with it. And they should take no payment that depends on what they find.

Two, plurality: several evaluation organisations across different risk areas, and permission to say where their conclusions differ from each other and from the company’s own staff.

Three, transparency. Methods, findings, the nature of the access. Labs should actively enable this, including by «limiting the scope of non-disclosure agreements». Evaluators should reach the board directly and unfiltered, publish findings, and accept only a time-limited redaction process.

Four, protection from retaliation: no revenge litigation, and funding that keeps flowing when the findings are awkward.

Five, access at the level of highly privileged employees. The same systems, data, tools and physical spaces as the senior internal people who run comparable risk assessments, plus candid one-on-one conversations with staff.

The letter points to the AEF-1 standard as one example of codified terms. And it is explicit that embedded evaluation complements broader external oversight rather than replacing it.

Who signed

Geoffrey Hinton and Stuart Russell head the list, alongside Arvind Narayanan and Yejin Choi. Then the evaluation business itself: Jacob Steinhardt and Sarah Schwettmann of Transluce, Adam Gleave of FAR.AI, Stella Biderman of EleutherAI, Miles Brundage, Rayan Krishnan of Vals AI, Charles Foster of METR.

Joy Buolamwini of the Algorithmic Justice League signed, so did Daniel Kokotajlo of the AI Futures Project and Vinh Nguyen, the NSA’s former chief AI officer. Everyone signed in a personal capacity; affiliations are listed for identification only.

Accenture fails at condition one

Read point one again and hold it next to Thursday’s announcement. Accenture is one of the largest IT services firms on the planet and does business with more or less everyone, Anthropic included. «No significant other commercial business» does not describe that relationship.

That is unlikely to be an accident. The letter landed a day after the announcement, and the signatories are precisely the organisations Anthropic did not pick first. METR is on the list of future partners, not on the first invoice.

The letter is useful anyway. It turns a warm word into a checklist. When Anthropic announces the next partnership six months from now, you can hold it against five points instead of a feeling.

Sources

AI SafetyAnthropicOpenAI