SDSignal Desk

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

Sep 21, 2026, 9:10 PM · MarkTechPost

Image: MarkTechPost

Grok 4.7 ships a larger base and longer RL run at Grok 4.6’s $2/$6 pricing — a price-performance shove, not a clean sweep of the leaderboard.

Why it matters

SpaceXAI released Grok 4.7 as its flagship for coding, agents, and knowledge work: new larger base model, longer reinforcement learning on hard multi-hour tasks, stronger self-verification claims, and native Grok Bot harness support — at the same $2 input / $6 output per million tokens as Grok 4.6.

It’s callable as grok-4.7 via the xAI API, Cursor, Grok Build, OpenRouter, Vercel, and Cloudflare, with a 500,000-token context window and knowledge cutoff of May 2026.

From the desk

Price freezes on a bigger model are the move. When Fable 5.1 Max sits at $10/$50 and GPT-5.6 Sol Max at $4/$20 on the launch table MarkTechPost reprints, holding $2/$6 is how you court Cursor defaults and API builders who feel every decimal.

Vendor-reported scores show Grok 4.7 beating Grok 4.6 across the board — Terminal-Bench 4.0 jumping 20.3% to 38.0%, EEBench to 64.0%, Harvey Legal Agent to 19.6%. It does not own the table: Fable leads four of seven listed benchmarks; GPT-5.6 Sol Max keeps DeepSWE v1.1 at 72.7%. Believe the direction, triangulate the absolutes.

Safety packaging matters for buyers who got burned by loose refusals. SpaceXAI calls the new safeguard stack its strongest yet, cites 62.4% on LatchBio’s biosafety benchmark, and says only 3.3% of risky dual-use prompts slipped through on its HackerBench v0.3 — while claiming not to block legitimate security work, with invite-only red-team access for partners.

Our take: useful competition at the mid-price tier is good for builders. A larger base plus longer RL on hard tasks is a coherent bet if latency stays honest — Grok 4.7 Fast doubles speed at double price, Cursor/Grok Build only. I’m watching independent Terminal-Bench and legal-agent replications, and whether the US regional endpoint’s 10% premium becomes a procurement checkbox.

Hype risk is treating the launch grid as physics. It’s a sales sheet. Ship evals on your own harness before you rewrite the stack around grok-4.7.

Context

MarkTechPost by Michal Sutter, September 21, 2026, summarizing SpaceXAI’s Grok 4.7 launch materials. All benchmark figures are vendor-reported as presented there.

Who feels it

API and Cursor users
Drop-in model id at unchanged $2/$6 pricing; Fast variant available only in Cursor/Grok Build at 2x price for 2x output speed.
Teams optimizing cost/performance
Strong paper on EEBench and Harvey legal agent relative to pricier frontier rows — validate on internal tasks.
Security and bio-risk reviewers
New refusal stack and HackerBench dual-use numbers invite third-party probing; invite-only cyber partner program expands defense research access.
Rival labs
Pressure to answer mid-tier price-performance without racing to the bottom on safety margins.

What to watch

  1. Third-party Terminal-Bench 4.0 and DeepSWE replications at matched reasoning effort.
  2. Default-model share inside Cursor and Grok Build after the switch.
  3. Independent audits of HackerBench dual-use pass rates versus external suites.
  4. Whether US-only endpoint demand grows among regulated customers.

Read the original

Continue at the source.

MarkTechPost

Companies: xAI