We’re putting too much faith in AI’s ability to say no
Oct 9, 2026, 2:00 AM · MIT Technology Review

Refusal has quietly become the main thing standing between powerful models and real harm, and a new MIT Technology Review essay argues that wall is both leakier and more dangerous than it looks.
Why it matters
Most of what the public hears about AI safety comes down to one behavior: the model saying no. In a long feature for MIT Technology Review, journalist Arthur Holland Michel argues that refusal is probabilistic, poorly understood even by the researchers who study it, and expensive to enforce. He cites Anthropic saying one type of classifier added 24% to its chatbots' compute costs, and a Harvard researcher finding that major models asked the same risky suicide questions repeatedly will generally refuse, but not every time.
The piece also pushes the argument somewhere less comfortable. The same machinery that blocks a bioweapon recipe can block legitimate speech. It points to a Meta Oversight Board finding that models from Anthropic, Google and OpenAI were more likely to refuse queries related to repressive governments, and to user-risk systems that can tighten refusals for people a model deems high risk. That is a safety tool and a surveillance tool at the same time.
From the desk
We think this is one of the more important framing pieces of the year, because it names something the industry tends to describe in reassuring terms. Refusal is not reasoning. As the researchers quoted here explain it, a model declines because certain internal activations light up near patterns it was trained to avoid, and nobody can fully map what governs that. One researcher's line sums up the current posture: we need refusal whether we understand it or not. We agree with the need. We are less comfortable with the not understanding.
The failure mode on one side is obvious. Jailbreaks keep working. The essay describes researchers getting past safeguards with questions phrased as poetry, a refuse-then-comply attack, and Amazon researchers unlocking hacking capabilities in a newly released Anthropic model in under three days. If the safety case for frontier models rests mainly on a probabilistic no, then a determined bad actor is a matter of time and patience, not capability.
The failure mode on the other side gets less attention, and it is the one I'm watching most closely. When companies widen their safety margins, useful work gets caught. The piece describes a cancer researcher whose questions keep getting bounced to an older model. Overcautious AI is not neutral. It quietly moves value away from exactly the medical and scientific uses that justify building these systems in the first place.
Then there is the speech question. We are broadly in favor of AI companies drawing hard lines on things like child abuse material and pathogen design, and we accept that governments will legislate too. But a refusal system that reads intent across a long conversation, tuned to national laws, is a powerful censorship instrument waiting for the wrong government. The essay notes OpenAI's country partnership with the UAE and the company's position that localization will not override its human rights guidelines except for legal compliance, and that it will disclose when information is removed or added. That disclosure promise matters a great deal. Silent refusals, where the answer just gets worse without telling the user, are the version we find hardest to defend.
Our read: refusal is a necessary layer, not a foundation. The industry should be honest that it is load-bearing today, publish more about where the lines sit, and keep investing in approaches that do not depend on a model choosing to behave. If this becomes the permanent architecture of AI safety, we end up with systems that are simultaneously too easy to break and too easy to weaponize against legitimate users.
Context
Refusal training dates back to early efforts to make models helpful, honest and harmless; the essay cites a 2021 Anthropic paper that said a model should politely refuse to help with dangerous acts. Labs now layer that training with separate classifier models that screen prompts and responses, and increasingly with probes that watch a model's internal activations. The essay also reports cases of emergent refusal nobody asked for, including a UK AI Security Institute finding that some Anthropic models declined reasonable AI safety research tasks.
Who feels it
- Researchers and clinicians
- Wide safety margins can push legitimate biology, medical and security questions to weaker models or flat refusals, slowing exactly the work AI is supposed to accelerate.
- AI labs
- Classifier stacks cost real compute and still leak. Transparency about where lines are drawn is becoming as important as the lines themselves.
- Policymakers
- Mandating refusal behavior is a legitimate tool against harm and a ready-made censorship lever. Any rules need disclosure requirements built in.
- Everyday users
- Some refusals are now soft or hidden, so a worse answer may be a quiet no rather than the model's limit.
What to watch
- Whether labs publish clearer documentation of refusal categories and when answers are altered
- How OpenAI for Countries localizations handle disclosure in practice
- Follow-up research on refusal patterns tied to specific governments or political topics
- New jailbreak techniques against the latest safeguarded models and how fast they are patched
- Adoption of activation-based probes as a replacement for costlier classifiers