Anthropic releases Claude Haiku 5.5 at $0.10 per million input tokens under 100K, about 75% cheaper than Haiku 4.5 on average; Sonnet 5.5 cache reads cut in half
The fastest and most efficient model in the Claude 5.5 family, built for high-volume, cost-sensitive work, and the first Haiku with an adjustable effort setting.
// Key points
- Pricing is tiered by prompt length: $0.10 per million input and $0.50 per million output tokens up to 100K tokens, $0.50 and $2.50 above that, versus $1 and $5 for Haiku 4.5. Anthropic notes the new tokenizer uses slightly more tokens per task, and it still comes out about 75% cheaper on average.
- Against Haiku 4.5: 72.4% vs 15.7% on OSWorld 2.1 (offline subset) and 39.2% vs 0.0% on Terminal-Bench 4.0. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, and Anthropic still recommends Sonnet 5.5 or Opus 5.5 for complex agentic coding.
- Sonnet 5.5 cache reads drop from $0.20 to $0.10 per million tokens starting today, which Anthropic says makes most agentic work about 20% cheaper; this week Max 5x and Max 20x subscribers start getting $100 and $200 in monthly API credits, and Team up to $500.
- The model ID is claude-haiku-5-5, available on the Claude Platform, AWS, Google Cloud and Microsoft Azure; the Python and TypeScript SDKs also add beta support for computer use and browser use.
Builder's takeI'm moving short-prompt calls, such as PandaClaws' classification, summaries and rewrites and AI Cloud Drive's document Q&A, onto Haiku 5.5 for a comparison run right away: under 100K tokens it costs a tenth of the old price, so a big model planning and Haiku doing the subtasks finally pencils out. The new tokenizer uses slightly more tokens, so compare cost per completed task rather than unit price; the cost table in my post on what it costs to keep three AI products online needs redoing too.