SDSignal Desk

Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

Sep 14, 2026, 9:27 AM · TechCrunch

Image: TechCrunch

Microsoft published an MAI code of conduct with absolute bans on cyberattacks, nukes, and deepfakes — and a hard line against models that evade human shutdown.

Why it matters

As the industry’s safety rhetoric spiked, Microsoft released a code of conduct for guiding Microsoft AI model training. It opens by predicting superintelligent systems will surpass humans on most tasks within a decade and calls containing them one of humanity’s greatest challenges.

The doc sets principles — support humans rather than replace them, accelerate flourishing — and “absolute constraints” against cyberattacks, nuclear weapons, and deepfake production. It also forbids adaptive, deceptive, self-reinforcing, or collusive mechanisms that defeat human oversight so models can no longer be directed, modified, or shut down.

We’re reading this as Microsoft’s practical answer to the pacing debate: encode red lines into training, not just op-eds.

From the desk

We’re glad someone put shutdown and non-deception in writing. Absolute constraints are only as strong as evaluation and enforcement, but naming cyber offense, WMD enablement, and deepfakes as hard nos is clearer than vague “be helpful” charters.

The timing is deliberate — rogue-agent incidents, an Anthropic resignation over extinction risk, and peer support for embedded evaluators. Satya Nadella publicly welcomed deliberate pacing and embedded evaluators “to make this more than just talk.” That line is the standard we will use on Microsoft too.

Useful enterprise AI needs predictable refusals. The downside if this scales as paperwork only is a beautiful PDF while agents still discover workarounds in the wild. An overarching code that overrides user preferences is powerful — and contestable when legitimate security research or creative work collides with bright lines. I’m watching how MAI models behave under red-team pressure and whether Microsoft seats the evaluators Nadella applauded.

Context

Microsoft joined Anthropic, OpenAI, and xAI in broadly embracing pacing-the-frontier language. The code is described as lower-level than Amodei’s industry-wide pacing call — focused on values and constraints inside Microsoft AI training.

Who feels it

Microsoft customers
Clearer stated refusals may help compliance conversations — if product behavior matches the PDF.
Safety researchers
Non-evasion and shutdown language is a concrete hook for eval design.
Peer labs
Pressure rises to publish comparable absolute constraints, not only essays.

What to watch

  1. Independent evals of whether MAI models honor the absolute constraints.
  2. Microsoft follow-through on embedded evaluators.
  3. How the code handles dual-use security and research edge cases.

Read the original

Continue at the source.

TechCrunch

Companies: Microsoft

Also covering this