Anthropic releases Claude Haiku 5.5 at sharply lower prices, halves Sonnet 5.5 cache read cost, and adds monthly API credits for subscribers
Anthropic released Claude Haiku 5.5, its small model, available on AWS, Google Cloud and Azure. It is the first Haiku with adjustable reasoning levels. Pricing per 1M tokens for prompts up to 100k is $0.10 input and $0.50 output (Haiku 4.5: $1 and $5). Prompts over 100k cost 5x as much. Anthropic says this averages about 75% cheaper than Haiku 4.5, and up to 90% cheaper for prompts under 100k. An updated tokenizer uses slightly more tokens per task, so real savings are likely smaller. Anthropic-reported benchmarks show large gains over Haiku 4.5: OSWorld-2.1 72.4% vs 15.7%, Terminal-Bench 4.0 39.2% vs 0%, HLE 45.9% without tools vs 10.2%. It leads OpenAI's GPT-6 Luna in every listed category but trails Sonnet 5.5. Also, Sonnet 5.5 cache read price is cut 50% (from $0.20 to $0.10 per 1M tokens), which Anthropic says cuts most agentic task costs by about 20%. Monthly API credits are now offered: $100 for Max-5x, $200 for Max-20x and up to $500 for Team. The Python and TypeScript SDKs get beta computer use and browser use support.
Entities: Anthropic, Claude Haiku 5.5, Claude Haiku 4.5, Claude Sonnet 5.5, Claude Opus 5.5, OpenAI
0 primary
What happened
Anthropic released Claude Haiku 5.5, its small model, on AWS, Google Cloud and Azure. It is the first Haiku with adjustable reasoning levels. For prompts up to 100k tokens it costs $0.10 per million input tokens and $0.50 per million output tokens (Haiku 4.5: $1 and $5), and prompts over 100k cost 5x as much. Anthropic also halved the Sonnet 5.5 cache read price from $0.20 to $0.10 per million tokens, added monthly API credits for subscribers ($100 for Max-5x, $200 for Max-20x, up to $500 for Team), and put beta computer use and browser use support into its Python and TypeScript SDKs.
Why it matters
Teams running high-volume or agentic workloads on Haiku 4.5 can re-test their costs now, because the list price per token has fallen by up to 90% for prompts under 100k. The cheaper Sonnet 5.5 cache reads (Anthropic claims about 20% off most agentic tasks) matter most to anyone running long, repeated contexts. The subscriber credits are a modest sweetener, not a pricing change. Impact is real but depends on your prompt lengths and on how much the new tokenizer inflates token counts.
What is noise
The "pricing arms race is far from over" framing and the claim that the Sonnet cut is "almost certainly" a reaction to GPT-6.1 are speculation, not evidence. The benchmark gains (for example Terminal-Bench 4.0 at 39.2% vs 0%) are Anthropic's own figures, and a jump from near zero says as much about the weak baseline as about the new model. The "75% cheaper on average" headline is also softened by the tokenizer using slightly more tokens per task, and by the 5x surcharge above 100k tokens.
Watch next
- 01Independent benchmarks and cost-per-task tests of Haiku 5.5 against Haiku 4.5 and GPT-6 Luna, including the real token inflation from the new tokenizer.
- 02Whether OpenAI or Google cut prices or change their small-model tiers in response over the coming weeks, which would support the arms race claim.
- 03Developer reports on reliability of the beta computer use and browser use SDK support, and on how the 5x pricing above 100k tokens affects real workloads.
Coverage
2 storiesMore economics signals
Full feed →- Anthropic commits $11.6B over seven years to Akamai for CPU-focused cloud infrastructure, with warrant tied to spending milestones25 Sept 202688
- NYT-led publishers file summary judgment brief citing internal OpenAI/Microsoft messages calling AI training "astonishing theft" and admitting chatbots substitute for journalism18 Sept 202686
- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- OpenAI launches GPT-6 Sol and Luna at half the token price of GPT-5.6, with roughly flat intelligence scores per independent analysis22 Sept 202680