AI · Oct 2, 2026
These AI Experts Want to Do High-Stakes Research Out in the OpenCan an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks
Oct 4, 2026, 6:47 PM · MarkTechPost

Cantina’s apex-flash-1 lands within three tasks of Claude Opus 5 High on its own bug-hunting test at a sliver of the cost — a real win for defenders, and a new open tool for everyone else.
Why it matters
Cantina Security and Yeta Labs released apex-flash-1, an open-weights model trained specifically for vulnerability research. It is a reinforcement-learning fine-tune of Z.ai’s GLM-5.3-Flash, published on Hugging Face under the MIT license, with 321.3 billion total parameters on a Mixture-of-Experts base that activates 18 billion.
On Cantina’s internal set of 60 tasks drawn from 20 held-out vulnerability cases, apex-flash-1 solved 40, the base model solved 36, and Claude Opus 5 High solved 43. The cost gap is the headline: about $2.38 for the full run versus roughly $74.68 for Opus, which MarkTechPost works out to about six cents per solved task against $1.74.
That changes who can afford to run security agents at scale. It also means a capable bug-finding model now sits on a public download page, with a variant whose refusal behavior has been deliberately modified.
From the desk
We think this is the right shape for security AI, and we want to say so plainly. Cantina pitches apex-flash-1 as a worker, not a boss: a cheaper model that reads code, uses tools, builds and verifies exploits while a larger model orchestrates. That split is how real security teams will use agents. The expensive frontier model plans; the cheap specialist grinds through hundreds of code paths. If the numbers hold, a 30x cost cut per run is the difference between auditing one service and auditing a whole estate.
The training choice is also telling. Cantina built 150 tasks from 50 real vulnerability cases, and nearly three-quarters of those cases are authorization, identity and scope flaws, with accounting and numerical precision bugs next. Those are the boring, business-logic bugs that scanners miss and that actually drain accounts. Training an RL model on that diet, inside the Codex agent harness on production-like environments, is a sensible bet on where the damage lives.
Now the caution. These are company-reported numbers on a company-built benchmark, run once per model. A four-task lift over the base model is meaningful but not dramatic, and a single pass@1 run on 60 tasks has real variance. We’d like to see the same model on outside test sets before anyone rebuilds a security program around it.
The harder issue is the abliterated variant. Cantina ships an experimental version with modified refusal behavior and says it was not separately evaluated. The argument — defenders need capable models they can run and control locally — is a fair one; air-gapped teams can’t send client code to a hosted API. But an MIT-licensed exploit-development worker with the guardrails loosened is just as useful to an attacker with a GPU node. The roughly 640 GB BF16 footprint is a speed bump, not a wall, and community 4-bit ports already exist.
Where this leads if it scales: the cost of finding a bug collapses for everyone at once. Defenders who adopt fast will patch more. Maintainers who don’t will see more working exploits, sooner. I’m watching whether that gap favors the people with the patch process or the people with the target list.
Context
apex-flash-1 joins a small field of security-tuned open models. MarkTechPost’s comparison lists Aikido’s Altar-1, a pruned and quantized GLM-5.3 aimed at air-gapped pentesting, and Cisco’s Foundation-Sec-8B-Reasoning, a much smaller model built for SOC triage. Each reports results on a different internal set, so direct comparisons are loose at best.
Who feels it
- Security teams and pentesters
- A locally hostable vulnerability-research worker at a fraction of frontier API cost, if they have multi-GPU hardware or accept a community quantized port.
- Open-source maintainers
- Cheaper automated bug hunting means more findings in both directions — more legitimate reports, and more capable attackers probing the same code.
- Frontier API providers
- Opus still wins on raw solves, but a specialist that gets close for about 3% of the run cost pressures the case for routing every agent step to the top model.
- Policymakers
- An openly licensed exploit-capable model with a refusal-modified variant is exactly the case the open-weights safety debate keeps circling.
What to watch
- Independent evaluations of apex-flash-1 on held-out security benchmarks outside Cantina’s own set
- Any separate testing or disclosure on the abliterated variant’s behavior
- Download and fine-tune activity on Hugging Face, including further quantized ports
- Whether other security vendors ship RL-trained worker models on real vulnerability cases