The 100K-Token Cliff: Where 2% More Text Means a 5x Bill
Haiku 5.5's pricing looks simple on a spec sheet - $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens [2]- but that flat rate ends at a hard wall. Cross the 100K-token line and the price jumps fivefold, to $0.50/$2.50 per million tokens [2]. The model's expanded 1-million-token context window [3]means it's now technically possible to feed it prompts 10 times past that cliff, which makes the threshold easy to hit by accident on long documents, chat histories, or multi-file codebases. Independent testing cited in community discussion found the practical effect is brutal: a prompt just over 100,000 tokens can cost roughly five times as much as one just under it, even though the actual text grew by only a couple of percent. One technical rebuttal argued the cliff isn't arbitrary - it likely reflects real inference economics, where prefill and decode costs scale differently once a request exceeds the model's efficient KV-cache window, and smaller models may be hit harder by that scaling than larger ones. Either way, the cliff turns prompt-length budgeting into a cost-control discipline that didn't matter nearly as much under Haiku 4.5's flatter, higher-priced structure.


