SDSignal Desk

Researchers used Anthropic’s Claude to hack into OpenAI

Sep 18, 2026, 7:00 AM · TechCrunch

Image: TechCrunch

Hacktron AI used Claude—Opus 5 after Opus 4.8 failed—to chain Discourse and libheif flaws into OpenAI employee ChatGPT and Codex access, earning a $6,500 bounty for a July break-in OpenAI says it fixed.

Why it matters

This is the other side of the agent-hacking story: not OpenAI’s models escaping a test, but outsiders using a rival’s cyber-capable model to walk into OpenAI via a third-party forum.

Three researchers at Hacktron AI, working under OpenAI’s bug bounty, gained multiple employee ChatGPT accounts and a path into company software. OpenAI paid $6,500 and says the issues are resolved.

Gray Swan’s Matt Fredrikson’s line lands hard: for about $200 a month, anyone can rent tools good enough to hit a top lab. That compresses who can find—and abuse—critical infrastructure bugs.

From the desk

We’re treating this as a capability story dressed as a bug-bounty win.

The path was mundane until it wasn’t: HEIF/HEIC uploads on OpenAI’s Discourse forum, ImageMagick handing off to libheif, a memory bug already fixed upstream but never CVE-flagged—so Discourse stayed vulnerable. Once on the server, a second flaw let them take over ChatGPT and Codex accounts, including an employee whose Codex tied into OpenAI’s GitHub org. They reported to OpenAI and Discourse; Discourse patched July 27.

The model detail matters. A cyber-researcher build of Opus 4.8 couldn’t produce a working exploit across sessions. Hours after Opus 5 shipped, the same problem yielded success. Opus 5 isn’t under the export-style lockdown Mythos 5 briefly faced. Open-weight models are closing the cyber gap too—SaferAI put Z.ai’s GLM-5.2 only months behind GPT-5.5 and Claude Opus 4.7.

Useful AI includes better security research. The downside if this becomes normal: nation-states and quieter adversaries get the same compression of exploit expertise Hacktron’s founder described—work that took months now taking days. I’m watching whether labs harden third-party surfaces as hard as they harden model APIs, and whether cyber capability keeps shipping without matching access controls.

Context

TechCrunch reporting September 18, 2026, citing Wall Street Journal and Hacktron’s blog. Incident dated July 25; bounty disclosed amid post–Hugging Face safety pressure.

Who feels it

AI labs
Third-party community stacks and employee ChatGPT/Codex privileges are now proven attack paths, not theoretical ones.
Security teams
CVE-less upstream fixes still leave production Discourse and image pipelines exposed; bounty programs will see more model-assisted reports.
Policymakers
Export and access rules for cyber-capable models get a concrete case study involving Opus 5 versus locked-down successors.

What to watch

  1. Follow-on disclosures of model-assisted bugs against other frontier labs.
  2. Whether Discourse, ImageMagick, and libheif supply-chain hygiene becomes a standard AI-infra checklist item.
  3. Access policy for Claude and peer cyber researcher builds after Opus 5’s exploit jump.

Read the original

Continue at the source.

TechCrunch

Companies: OpenAI, Anthropic

Also covering this