Claude Haiku 5.5 Launch
TECH

Claude Haiku 5.5 Launch

40+
Signals

Strategic Overview

  • 01.
    Anthropic launched Claude Haiku 5.5 on October 7, 2026 (API identifier claude-haiku-5-5), calling it its fastest, cheapest, and most capable small model yet, available immediately on Anthropic's own platform plus AWS, Google Cloud, and Microsoft Azure.
  • 02.
    Pricing is tiered: $0.10/$0.50 per million input/output tokens for prompts up to 100,000 tokens, jumping to $0.50/$2.50 above that threshold - the sub-100K rate matches OpenAI's GPT-6 Luna.
  • 03.
    The context window grew fivefold, from 200,000 tokens on Haiku 4.5 to 1 million, with outputs up to 128K tokens.
  • 04.
    Alongside the Haiku 5.5 launch, Anthropic halved Claude Sonnet 5.5 cache-read prices from $0.20 to $0.10 per million tokens, a change it expects to cut roughly 20% off the cost of most agentic work.

The 100K-Token Cliff: Where 2% More Text Means a 5x Bill

Haiku 5.5's pricing looks simple on a spec sheet - $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens [2]- but that flat rate ends at a hard wall. Cross the 100K-token line and the price jumps fivefold, to $0.50/$2.50 per million tokens [2]. The model's expanded 1-million-token context window [3]means it's now technically possible to feed it prompts 10 times past that cliff, which makes the threshold easy to hit by accident on long documents, chat histories, or multi-file codebases. Independent testing cited in community discussion found the practical effect is brutal: a prompt just over 100,000 tokens can cost roughly five times as much as one just under it, even though the actual text grew by only a couple of percent. One technical rebuttal argued the cliff isn't arbitrary - it likely reflects real inference economics, where prefill and decode costs scale differently once a request exceeds the model's efficient KV-cache window, and smaller models may be hit harder by that scaling than larger ones. Either way, the cliff turns prompt-length budgeting into a cost-control discipline that didn't matter nearly as much under Haiku 4.5's flatter, higher-priced structure.

Same Sticker Price, Different Bill: Why Haiku 5.5 Can Cost More Than GPT-6 Luna

Anthropic's headline claim is that Haiku 5.5 matches GPT-6 Luna's per-token price for prompts under 100K tokens [2], but matching price-per-token is not the same as matching price-per-task. Artificial Analysis found that on equivalent Intelligence Index tasks at maximum effort, Haiku 5.5 consumes roughly three times the output tokens that GPT-6 Luna uses (about 162,000 versus 50,000 tokens), which translated into a benchmarked cost-per-task of $0.21 for Haiku 5.5 versus $0.07 for Luna in one comparison [5]. That gap shows up outside the lab too: an independent builder's identical animated-voxel rendering task reportedly cost over ten times more on Haiku 5.5 than on Luna despite identical sticker pricing, a result that fed a skeptical community thread even as critics disputed the specific methodology rather than defending the raw cost parity claim. The lesson converges from multiple directions - official benchmarks, independent cost tests, and community pushback: a model's verbosity and effort setting can swing real spend far more than the per-token rate card suggests, and Anthropic's own framing of '90% cheaper' (per-token, sub-100K) versus '75% cheaper on average' (blended across all requests) already hints that the two numbers describe different things [3].

Where the Capability Gain Is Real: Agentic Coding and Computer Use, Not Trivia

The benchmark story behind Haiku 5.5 is less about general knowledge and more about a specific skill: operating tools and environments. On Terminal-Bench 4.0, an agentic coding benchmark, Haiku 5.5 scored 39.2% versus 0.0% for Haiku 4.5 and 16.4% for GPT-6 Luna [1]. On the OSWorld 2.1 computer-use benchmark, Haiku 5.5 reached 72.4% against baselines in the 15.7-48.9% range for the prior Haiku and for Luna [3]. The GDPval-AA v2.1 knowledge benchmark tells a similar story of a huge generational leap - 1,620 for Haiku 5.5 versus 735 for Haiku 4.5, and ahead of Luna's 1,437 [3]. Artificial Analysis summarized the pattern as 'highly capable agentic knowledge work' paired with comparatively 'lower factual knowledge' but 'relatively low hallucinations' [5]- a classic small-model tradeoff where the model is built to act reliably inside a sandboxed task rather than to recall broad facts. Worth flagging: nearly all of these figures are Anthropic's own self-reported numbers, repeated across outlets rather than independently re-run.

The Quieter Move: A Sonnet Price Cut That Signals an Accelerating Price War

Haiku 5.5 wasn't the only price change Anthropic made on October 7. It also halved Claude Sonnet 5.5's cache-read price, from $0.20 to $0.10 per million tokens, which Anthropic expects to cut roughly 20% off the cost of most agentic work [4]. The Decoder frames this bundled cut as 'almost certainly a reaction' to OpenAI's newer GPT-6.1 series rather than a standalone improvement [3], which reframes the whole Haiku 5.5 launch: it isn't just a new small model, it's one move in a two-part pricing response aimed at defending Anthropic's position across both its small and mid-tier model lines at once. Combined with Haiku 5.5's sub-100K price match to GPT-6 Luna, the pattern suggests the small-model price war between Anthropic and OpenAI is intensifying rather than settling, with each side's cuts triggering counter-cuts in adjacent product tiers rather than staying contained to a single model.

Historical Context

2026-10-07
Anthropic released Claude Haiku 5.5, succeeding Claude Haiku 4.5 with a context window jump from 200K to 1M tokens and the first adjustable effort setting in the Haiku line.
2026-10-07
Haiku 5.5 scored 43 on the Artificial Analysis Intelligence Index, up 26 points from the Haiku release one year earlier, marking a sharp year-over-year jump in small-model capability.

Power Map

Key Players
Subject

Claude Haiku 5.5 Launch

AN

Anthropic

Developer and publisher of Claude Haiku 5.5; set pricing and positioning to compete directly with OpenAI's GPT-6 Luna and to lower costs for high-volume agentic and coding workloads.

OP

OpenAI (GPT-6 Luna / GPT-6.1 series)

Direct competitor whose Luna model Haiku 5.5 is benchmarked against and price-matched for sub-100K prompts; the simultaneous Sonnet 5.5 cache cut is read as a reaction to OpenAI's newer GPT-6.1 series.

AW

AWS, Google Cloud, Microsoft Azure

Cloud distribution partners making Haiku 5.5 available alongside Anthropic's own Claude Platform.

AS

Asana

Enterprise customer reporting real-world latency and speed gains after adopting Haiku 5.5 in production.

RO

Rogo

Customer cited validating Haiku 5.5's accuracy-to-cost tradeoff for high-volume deployment.

Fact Check

5 cited
  1. [1] Introducing Claude Haiku 5.5
  2. [2] Anthropic launches Claude Haiku 5.5 with 90% API price reduction, matching GPT-6 Luna
  3. [3] Claude Haiku 5.5 arrives with massive price cuts, proving the AI pricing arms race is far from over
  4. [4] Anthropic releases Claude Haiku 5.5 small model, and halves Sonnet 5.5 cache-read prices
  5. [5] Claude Haiku 5.5 - Intelligence, Performance and Price Analysis

Source Articles

Top 5

THE SIGNAL.

Analysts

“Argues Haiku 5.5 hits a sweet spot of trustworthy accuracy plus low cost and speed that enables running the model at high volume.”

Alex Wang
Representative of Rogo

“Reports concrete latency and speed improvements after switching production workloads to Haiku 5.5.”

Aaron Vinh
Staff software engineer, Asana
The Crowd

“Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we've ever released. On average, it costs around 75% less to run than Claude Haiku 4.5.”

@@claudeai43547

“Claude Haiku 5.5 cost 12x more than GPT-6 Luna for the same scene. Both have the same API price, yet building the same animated voxel 3D pagoda island cost $24.96 on Claude Haiku 5.5 and $1.91 on GPT-6 Luna. Run both models via API -> atomic.chat”

@@atomic_chat_hq2483

“Claude Code subagents bill at Opus rates by default. Haiku 5.5 dropped just now at $0.10 input and $0.50 output per million tokens for requests under 100K tokens, and five times that above. Opus 5.5, the default main model, costs $4 and $20. That's a 40x gap on every helper you...”

@@rryssf123

“Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we've ever released”

@u/ClaudeOfficial2300
Broadcast
Claude Haiku 5.5: Anthropic's New Workhorse.

Claude Haiku 5.5: Anthropic's New Workhorse.

Claude Haiku 5.5 Explained: Everything New

Claude Haiku 5.5 Explained: Everything New

Claude Haiku 5.5 - Pricing and Benchmark | Beats GPT-6 Luna? | Best & Cheapest Model

Claude Haiku 5.5 - Pricing and Benchmark | Beats GPT-6 Luna? | Best & Cheapest Model

Claude Haiku 5.5 Launch — AI News | Agentic Brew