SDSignal Desk

Anthropic Releases Claude Haiku 5.5: A Small Model With 1M Context Priced at $0.10 per Million Input Tokens

Oct 7, 2026, 1:38 PM · MarkTechPost

Image: MarkTechPost

Haiku 5.5 makes a capable model with a million-token memory nearly free for everyday work. Cheap intelligence at this price changes what gets automated, and how much.

Why it matters

Anthropic released Claude Haiku 5.5, which it calls its cheapest and fastest small model yet. It keeps a 1 million token context window, up to 128,000 output tokens, and costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Above that, rates rise.

Anthropic estimates it runs about 75% cheaper on average than Haiku 4.5 after accounting for a new tokenizer that counts the same text as roughly 30% more tokens. It is generally available on Anthropic's API, Amazon Bedrock, Google Cloud and Microsoft Foundry. And it matches OpenAI's GPT-6 Luna on short-prompt list price, which tells you where the small-model price war has landed.

From the desk

We think this is good news for useful AI, mostly. Small, cheap models are the workhorses. They summarize documents, sort tickets, pull a number out of a filing and hand it to a bigger model. When that layer gets dramatically cheaper and better, a lot of tedious work becomes economical to automate, and smaller companies get access to capabilities that used to need a serious budget.

The reported jump in capability is large. Anthropic's own numbers put Haiku 5.5 at 72.4% on an offline subset of OSWorld 2.1 for computer use, against 15.7% for Haiku 4.5. That is a different class of model, if it holds up. Anthropic is also honest about positioning: it still recommends Sonnet 5.5 and Opus 5.5 for complex agentic coding, and pitches Haiku as the subagent working under them.

Now the fine print, because it matters. The headline price applies under 100,000 tokens. Above that, the million-token window costs five times as much per token, and GPT-6 Luna's higher tier kicks in much later, so for long prompts Luna is cheaper on list price. The new tokenizer quietly raises token counts, and non-default sampling settings now return errors. Teams migrating should run their own numbers, not the headline.

The bigger downside is what near-free intelligence does at scale. When a computer-use-capable model costs this little, it becomes cheap to run thousands of agents clicking through websites, filling forms and generating content. That powers great customer support and automation. It also powers spam, scraping and fraud at volumes we have not seen. I'm watching whether usage controls keep pace with price cuts.

Our read: a strong release that makes the agent stack cheaper from the bottom up. All benchmark figures are Anthropic-reported, so independent testing is the next thing to wait for.

Context

Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, with adaptive thinking on by default. It accepts text and images and has a June 2026 knowledge cutoff. Anthropic cited early users including Rogo, which uses it as a subagent pulling data from filings, and AlphaSense, which tested it on a feature handling about 8 million calls a week.

Who feels it

Developers
Much cheaper subagents and summarization, but migration requires checking tokenizer changes and sampling parameter limits.
Enterprises
High-volume document and support workloads get cheaper; long-context jobs above 100,000 tokens need separate cost modeling.
Competing labs
Price parity with GPT-6 Luna keeps pressure on small-model pricing across the industry.
Trust and safety teams
Cheap computer-use agents lower the cost of automated abuse as well as automation.

What to watch

  1. Independent benchmarks of Haiku 5.5 on computer use and coding
  2. Real-world cost comparisons after the tokenizer change
  3. OpenAI and Google responses on small-model pricing
  4. Abuse patterns tied to cheap computer-use agents

Read the original

Continue at the source.

MarkTechPost

Companies: Anthropic