Anthropic Launches Claude Haiku 5.5 with 90% Price Cut from Previous Generation
nashnova research
Anthropic launched Claude Haiku 5.5 at a headline ~90% price cut from Haiku 4.5, though real-world savings average about 75% — a new tokenizer consuming ~30% more tokens and only a 50% discount on long-context requests eat into the gap. Benchmarks leap past same-tier rivals, but complex coding still falls well short of Sonnet 5.5. The model is positioned as a high-throughput execution layer.
The headline says 90% cheaper — what do you actually save?
Short-context pricing: $0.10 per million input tokens, $0.50 output — one-tenth of Haiku 4.5 and one-twentieth of Sonnet 5.5.
But Anthropic's own estimate puts real savings at about 75% on equivalent tasks. This means → two hidden costs erode the discount: ① requests over 100K tokens get only a 50% cut; ② the new tokenizer makes the same text consume roughly 30% more tokens.
In plain terms = if your calls stay under 100K tokens, savings track close to the headline; once context gets long, the discount shrinks fast.
How does it stack up against GPT-6 Luna?
Under 100K tokens, Haiku 5.5 and GPT-6 Luna are priced identically.
For long context, Luna is cheaper: it doesn't raise prices until roughly 270K tokens, then charges $0.20 input / $0.75 output. Haiku 5.5 jumps at 100K to $0.50 input / $2.50 output.
This means → for workloads centered on long documents, Luna holds a clear cost edge; for short calls, the two are interchangeable on price.
Benchmarks leap ahead of its tier — but where's the ceiling?
OSWorld 2.1 — testing an AI agent operating a real computer through multi-step tasks: Haiku 5.5 Max hits 72.4%, same-price Luna scores 48.9%, previous-gen Haiku 4.5 just 15.7%. A generational jump.
Terminal-Bench 4.0 (coding): Haiku 5.5 scores 39.2%, Luna 16.4%, but Sonnet 5.5 reaches 70.6%. This reflects that complex coding remains the sharpest divide between large and small models.
GDPval-AA (knowledge work): Haiku 5.5 hits 1,620, Luna 1,437, Sonnet 5.5 1,840. Anthropic explicitly recommends Sonnet 5.5 or Opus 5.5 for complex agentic coding.
What is it built for?
Anthropic positions Haiku 5.5 as the execution-layer model — high volume, fast, cheap. Target use cases: summarization, context compression, database queries, content classification, real-time customer support.
In plain terms = it's not the thinker; it's the fast hand on the assembly line — simple tasks, run quickly and cheaply.
Real-world deployments: Cognition's Devin pairs Opus 5.5 as the lead model with Haiku 5.5 as the sub-agent, scoring 66.2% on FrontierCode — higher than either model alone — while cutting cost and latency. Financial AI firm Rogo uses it to extract segment revenue data from 10-K filings at scale.
What should developers watch when migrating?
Five key API changes: `budget_tokens` must switch to adaptive thinking plus effort tiers; `temperature`, `top_p`, `top_k` parameters must be removed; assistant-message prefilling is dropped; the computer-use tool must upgrade to `computer_toolset_20260801`; adaptive thinking is on by default, so the first content block may be a thinking block — parsing logic must filter on the `type` field.
The new tokenizer inflates token counts by roughly 30% for the same text, so `max_tokens` settings and cost estimates both need recalculating.
This means → you can't just swap the model name and ship; the interface layer and the cost budget both need a fresh pass.
What else shipped alongside?
Sonnet 5.5 cache-read pricing drops from $0.20 to $0.10 per million tokens; Anthropic estimates most agent workloads save about 20% as a result.
Max and Team subscribers can now claim monthly API credits starting this week: Max 5x gets $100, Max 20x gets $200, Team up to $500 (shared across the team).
Haiku 5.5 is live on the Claude client, Claude Code, AWS, Google Cloud, Azure, Cursor, and OpenRouter. Context window: 1 million tokens; max single output: 128K tokens. Only Fable 5.5 remains unreleased in Anthropic's model lineup.
市场有风险,内容仅供研究参考,不构成投资建议。
