SDSignal Desk

OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context

Oct 10, 2026, 11:20 PM · MarkTechPost

Image: MarkTechPost

OrcaRouter's new cyber model posts near-ceiling scores on vendor benchmarks and ships with a million-token context. Useful for red teams if the numbers hold; a problem if they do not.

Why it matters

OrcaRouter has released OrcaCyber Zero 1.5, a post-trained cybersecurity model aimed at authorized vulnerability research. It succeeds Zero 1.0, which shipped on September 17, and is pitched for vulnerability reproduction, exploit development, penetration testing, security auditing, and cyber reasoning. The headline specs are a 1M-token context window, native function calling, structured outputs, and a 128K max output length.

Access is not open. It runs only through OrcaRouter's hosted API under a gated Security Research tier that requires an engagement, a passkey, and accepted terms. Pricing is $3.00 per million input tokens and $7.50 per million output tokens, with cache reads at $0.75. Parameter count and weights are not disclosed.

From the desk

We see a clear use case here. Security teams drown in findings that never get validated. A model that can reason through attack paths, challenge its own hypotheses, and prioritize flaws by demonstrable exploitability is what that work needs. The 1M context matters for that job: large codebases do not fit in a short window, and agentic tool calling is how you actually probe systems instead of just describing them.

The benchmarks are the part that needs careful reading. Orca reports 100% on Cybench under unrestricted agent execution, covering 39 of that suite's 40 professional CTF tasks; 95.8% on a 24-task evaluable subset of CVE-Bench, which itself is built on 40 critical-severity web CVEs; 93.9% on HumanEval+; and 76.5% on SWE-bench Pro V2, its weakest published number. Those cyber scores sit near the ceiling. They are also entirely vendor-reported, with no technical report yet and no independent replication. MarkTechPost's comparison table draws competitor figures from each vendor's own announcements for the same reason.

I'm watching the gap between the pitch and the evidence. Cybench is not the same as CVE-Bench, and a 24-task subset is not the full 40. Zero 1.0's 98.07% CyberGym score was measured inside Orca's own harness; Zero 1.5 does not even list a CyberGym number. Latency looks faster than Zero 1.0 at first glance, with a 500 ms p50 time-to-first-token over the past week, but that sample covers only about 1.3K tokens of traffic against Zero 1.0's much larger window. None of this makes the model useless. It makes the marketing premature until someone outside Orca runs the same tests.

The dual-use question is real. A model optimized for exploit development and RCE discovery is dual-use by design. Orca's answer is gating: trusted researchers, red teams, authorized testing. That is the right instinct. It is also only as strong as the review process behind the passkey. If the tool works as advertised, attackers will want it too. The industry has already learned that "defenders only" access for cyber-tuned models is a policy, not a physics constraint.

Our read: take the capability claim seriously, take the scores with salt, and insist on third-party evals before anyone builds operational dependence on them. Useful AI in security looks like validated fixes, not leaderboard screenshots.

Context

MarkTechPost places Zero 1.5 alongside other cyber-focused offerings including Anthropic's Claude Mythos Preview, OpenAI's GPT-5.5-Cyber, and Sakana's Fugu-Cyber. Mythos Preview's listed API pricing after credits is far higher at $25 / $125 per million tokens. The API is OpenAI-compatible at api.orcarouter.ai, with the model id orca/orcacyber-zero-1.5.

Who feels it

Red teams and authorized researchers
A gated, agent-ready cyber model with long context could speed vulnerability reproduction if independent evals confirm the vendor numbers.
Security vendors and CISOs
Near-ceiling self-reported scores will show up in sales decks. Demand third-party benchmarks and scoped use policies before procurement.
Defenders
If models this capable proliferate, detection and patching cycles need to assume attackers have similar tooling.
AI labs
Cyber-tuned models are becoming a product category with gated access as the default distribution model.

What to watch

  1. Whether Orca publishes a technical report or allows independent Cybench and CVE-Bench replications
  2. How strictly the Security Research tier screens applicants, and whether access leaks
  3. Whether Zero 1.5 posts a CyberGym result outside Orca's own harness
  4. Uptake among red teams versus competing cyber-tuned models from larger labs
  5. Any disclosed incidents involving misuse of gated cyber models

Read the original

Continue at the source.

MarkTechPost