AI · Sep 2, 2026
Google releases Gemini 3.8 Flash, its third Flash model in six weeksGoogle says its new Gemini 3.8 Flash model ‘works harder’ but might cost more
Sep 2, 2026, 1:11 PM · The Verge

The Verge's Gemini 3.8 story is the bill: same $0.75 / $3.75 sticker as 3.7, about 40 percent more expensive per task because it talks more.
Why it matters
Stevie Bonifield reports Gemini 3.8 Flash launched a few weeks after 3.7. Google says it 'works harder' by doing more reasoning steps and calling tools iteratively. Introductory pricing matches 3.7: $0.75 per million input tokens and $3.75 per million output. Google warns the model might use more tokens, especially at higher effort, and says developers can keep 3.7 to minimize usage.
Artificial Analysis called 3.8 Flash the cheapest it has measured at this intelligence level, but said cost is up about 40 percent from 3.7 despite unchanged per-token pricing, driven by a 30 percent increase in output tokens per task and more turns on agentic evaluations. Aigora.ai CEO John Ennis compared it to Anthropic, claiming 'Opus 5 coding quality but at a fraction of the cost and super fast.' Google says 3.8 beats 3.7 and other frontier models on DeepSWE v1.1, including Anthropic's Fable 5, and also leads Vals Finance Agent V2 and Harvey's Legal Agent benchmark. It ships with CBRN and cyber-offense safeguards. 3.8 Flash Cyber launches into the Fairwind Program, limited to governments and trusted partners — 650 members including CrowdStrike and the Center for Internet Security — with Google's CodeMender agent. 3.8 Flash is out now for Google AI Pro and Ultra consumers plus developers and enterprise.
The Signal Desk read
Bonifield's piece is the one that treats token pricing as a trick. Artificial Analysis did the arithmetic Google buried: same rates, 30 percent more output tokens, more agentic turns, ~40 percent more dollars per task. 'Cheapest at this intelligence' can be true at the same time as 'more expensive than last month's Flash.' That is the whole story for anyone on a metered API.
Ennis's Opus 5 line is a tweet, not a benchmark. Leave it as color. DeepSWE versus Fable 5 is the comparison Google wants, timed against Anthropic's own cheaper-cached-data pitch the same week. Flash versus a named Anthropic coding model is the market Google thinks it can take on unit price — until effort-level token burn eats the discount.
Signal Desk's read: 'works harder' is a positioning gift and a finance problem. Users who liked Flash because it was cheap will notice the bill before they notice DeepSWE. Google even left 3.7 up as the efficiency SKU, which is an admission. Fairwind's 650-member list, CrowdStrike and CIS named, is how Cyber gets a constituency without a public download. CodeMender bundled in is the upsell: not just a model, a patch loop for infrastructure.
If Artificial Analysis's 40 percent holds up outside their harness, Google shipped a price increase and called it a model launch.
Context
Flash has become Google's weekly-ish coding and agents SKU. Token prices across labs have been falling to keep skittish buyers in. 3.8 tests whether 'introductory' rates survive contact with a model instructed to think longer.
Who feels it
- Metered API users
- Budget in 30–40 percent more tokens per hard task, or pin 3.7. Sticker price is not the bill.
- Anthropic customers
- Google is aiming 3.8 at Fable/Opus coding. Run the same evals; ignore launch tweets.
- Governments and CISOs
- Fairwind plus CodeMender is the gated cyber bundle. Membership is the product.
What to watch
- Whether production traces confirm Artificial Analysis's 30 percent output-token jump.
- How many teams actually stay on 3.7 for cost.
- Fairwind expanding past the 650, or staying a club.
Companies: Google