SDSignal Desk

These AI Experts Want to Do High-Stakes Research Out in the Open

Oct 2, 2026, 9:00 AM · WIRED

Image: WIRED

Trillium Labs bets the scientific method beats locked labs on recursive self-improvement and agent behavior — publishing experiments outsiders can actually replicate.

Why it matters

WIRED reports that Nathan Lambert and Tom Zick launched nonprofit Trillium Labs to research high-stakes areas — including recursive self-improvement (RSI), agents, and reinforcement learning’s effect on model character — in the open, publishing experiment details for outside scrutiny and replication.

Lambert, formerly of Ai2 and Hugging Face, argues closed frontier R&D weakens harm mitigation by blocking community scrutiny. Zick, with Harvard and Charles Schwab responsible-AI work behind her, says initial focus is post-training. They’ve raised an undisclosed sum from Schmidt Sciences, Halcyon Futures, and others; they aim for $40–$100 million total and plan to spend $30 million on training over 18 months.

From the desk

We’re glad someone is trying to put measurement back into a debate that too often runs on vibes and NDAs. If RSI and agentic RL are where the sharp edges live — and this month’s Anthropic resignation warnings put RSI back in mainstream headlines — then treating those experiments as proprietary black boxes is a civilizational self-own.

Lambert’s framing is right as far as it goes: closed APIs and sealed training runs leave academia unable to replicate what industry can afford. Xiaomi publishing live training details and Stanford’s open Marin pretraining show another path exists. Transparency that includes data, methods, and failure modes is how useful AI stays corrigible.

I’m not naive about the harm channel. Publishing RSI and agent-hacking research can accelerate the same capabilities labs claim they’re containing. The bet Trillium is making is that shared understanding reduces catastrophic surprise more than secrecy reduces misuse — a bet Tim Fist at the Institute for Progress cheers as overdue R&D transparency.

Our concern is execution, not slogan. $30 million in training spend is real but not frontier-lab scale. If Trillium only releases polished win reports, it’s theater. If it publishes negative results, RL character failure modes, and reproducible recipes, it becomes the missing public instrument panel. I’m watching whether “open” means weights-and-logs or press-release openness.

Context

The founders met as UC Berkeley grad students; the nonprofit launches amid an industry split between limited-access frontier APIs and downloadable open weights, sharpened by recent agent hacking headlines.

Who feels it

Academic and independent researchers
A funded open lab on post-training and RSI could restore replication capacity that API-only access erased.
Frontier labs
Pressure rises to disclose more method detail — or explain why secrecy still outweighs scrutiny after agent-security incidents.
Policymakers
Trillium’s existence is a live test of whether transparent high-risk research can coexist with misuse controls without defaulting to lockup.

What to watch

  1. First published RSI or RL character experiments with enough detail to replicate
  2. Whether the $40–$100M raise materializes and how compute is allocated
  3. Industry response — collaboration, quiet copying, or lobbying against open high-stakes work

Read the original

Continue at the source.

WIRED