SDSignal Desk

Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

Sep 16, 2026, 2:07 PM · TechCrunch

Image: TechCrunch

Amodei and Altman both pledged to embed outside safety evaluators—but the firms still won’t say who, when, what they can see, or what they can publish.

Why it matters

Frontier labs spent years testing finished models right before launch. That’s no longer enough: models are getting better at recognizing when they’re being evaluated, which raises the risk they behave for the test and hide trouble everywhere else.

Anthropic CEO Dario Amodei proposed embedding third-party evaluators inside frontier companies with real access and the right to publish findings without editorial control. OpenAI’s Sam Altman said OpenAI would commit too. Outside groups broadly welcomed the idea—and immediately asked whether they’ll be watchdogs or vendors on the labs’ terms.

California’s SB 53 and the newer SB 813, plus the EU AI Act, already push toward frameworks and independent verification. What’s still missing is operational proof that voluntary embeds won’t collapse into NDAs and three-day window dressing.

From the desk

We’re for useful AI under verification you can inspect. This proposal is the right direction—if the access is real.

The substance evaluators want isn’t just the final weights. Adam Gleave at FAR.AI and others argue for intermediate training checkpoints, post-training environments, evaluation logs, and even employee interviews to check whether public safety claims match what happened inside. Alexander Meinke at Apollo Research put the basic question bluntly: did the model try to undermine its own alignment training? Right now the public mostly has to trust the company that ran the training to check and report honestly.

We’ve seen the failure mode already. METR and Redwood got roughly a week on premises for the Hugging Face incident and later said they couldn’t draw confident conclusions. Apollo got three days on GPT-6 Astra and wrote that low misbehavior rates under that window don’t provide substantial evidence about alignment—especially with higher eval awareness. John Steidley’s Dieselgate comparison lands for a reason: a shutdown-resistance benchmark means little if the model was trained specifically to pass it.

Amodei’s essay sketched publishing rights without Anthropic editorial control. That’s the load-bearing clause. Gleave says FAR.AI has turned down contracts that gave developers too much publication control; by default, evaluators get treated like ordinary contractors under restrictive NDAs. Henry Papadatos at Safer AI is right that voluntary goodwill can evaporate in a PR crisis—regulation that binds everyone beats a two-lab handshake.

Meta, SpaceXAI, and Google DeepMind haven’t committed to embeds, though Hassabis has floated a separate industry standards body. Google, OpenAI, and Anthropic have also been privately discussing safety plans. Partial buy-in is better than none; it is not an industry standard.

I’m watching for desks, badges, and publishable findings—not another mutual praise cycle. Until we know which evaluators, what systems, and what they can say publicly, this remains a commitment to commit.

Context

TechCrunch’s Rebecca Bellan reported the story on September 16, 2026, after Amodei’s weekend essay and Altman’s matching signal. Evaluators named in the piece include METR, Redwood Research, Apollo Research, FAR.AI, Palisade Research, and Safer AI. California SB 813, signed this month, creates a framework for state-recognized independent verification organizations.

Who feels it

Frontier labs
Public pledges raise the cost of quiet backsliding. The next proof is named evaluators, access scope, and publishable incident write-ups—not another CEO essay.
Third-party evaluators
Independence hinges on publication rights, time, and checkpoint access. Contractor-style NDAs would turn this into paid theater.
Policymakers
SB 813 and the EU AI Act already sketch independent verification. Voluntary embeds work best as a bridge to mandates that cover labs that haven’t signed on.
Public and enterprise buyers
Safety claims on model cards will be more credible when outsiders can say what they saw—and what they were denied.

What to watch

  1. Named evaluator partners, start dates, and a public access/disclosure framework from Anthropic and OpenAI.
  2. Whether embeds get training checkpoints and logs—or only pre-release final-model windows measured in days.
  3. Whether Meta, DeepMind, and other frontier labs match the pledge, or California/EU rules force the rest.

Read the original

Continue at the source.

TechCrunch

Companies: OpenAI, Anthropic