Anthropic released Claude Haiku 5.5, its fastest small model, priced at about a quarter of the running cost of Haiku 4.5. For prompts up to 100,000 tokens the new rates are $0.10 per million input tokens and $0.50 per million output tokens. Cache reads on Claude Sonnet 5.5 fall from $0.20 to $0.10 per million tokens. The move matters because it lowers the unit economics of high-volume agents, support automation and classification.
Pricing and positioning of Haiku 5.5
Haiku 5.5 completes a three-model 5.5 generation two weeks after Opus 5.5 launched on Sept. 22, with Sonnet 5.5 having launched on Sept. 28. Haiku 4.5, released last October, costs $1 per million input tokens and $5 per million output tokens. Anthropic said prompts up to 100,000 tokens covered about 90% of requests on the older model, and in that band Haiku 5.5 is 90% cheaper on input and output alike. Longer prompts receive a 50% discount, while the quoted 75% average saving accounts for a new tokenizer that uses slightly more tokens per task.
Anthropic positions Haiku 5.5 for repetitive work such as high-volume summaries and classification. Coding teams can deploy it as a subagent to which Opus 5.5 or Sonnet 5.5 delegates smaller tasks. The company describes it as the fastest Anthropic model at standard speed and recommends it for live customer support and browser automation. It is also the first Haiku model with an adjustable effort setting, which allows cost to be tuned against output quality for specific workloads.
The pricing matches OpenAI GPT-6 Luna, launched last month at the same $0.10 and $0.50 rates. Anthropic published benchmarks show Haiku 5.5 ahead of Luna on all six tests where both have scores. On the OSWorld 2.1 offline subset for long multistep computer operation, Haiku 5.5 scored 72.4% against 48.9% for Luna. On Terminal-Bench 4.0 for agentic coding, the scores were 39.2% against 16.4%. For complex agentic coding, Anthropic still directs customers to Sonnet 5.5 and Opus 5.5.
What lower agent costs mean for business
For companies running agents at scale, the combination of cheaper input, cheaper output and cheaper cache reads changes which tasks can be automated profitably. Cached tokens account for a large share of consumption, and Anthropic expects the Sonnet 5.5 cut to reduce the cost of most agentic work on that model by about 20%. A small firm can now process larger volumes of tickets, documents or product data without moving to a weaker open model. A large organization can split work across tiers, keeping Sonnet or Opus for planning and Haiku for execution, summaries and tool calls.
Early testing points to speed gains alongside price, but buyers should verify them on their own flows. Asana tested Haiku 5.5 in its AI Teammates agent evaluation suite and reported task completion latency more than 30% lower than its current model, with inference per agent turn up to 2.5 times faster. Safety testing showed far fewer instances of misaligned behavior than Haiku 4.5, while cybersecurity rules allow more defensive work than Sonnet 5.5 permits. Penetration testing remains blocked along with other attacker-oriented techniques, with wider access available only through the Cyber Verification Program expanded on Tuesday.
Availability and subscription terms will show whether the savings reach production quickly. Haiku 5.5 is available now on the Claude Platform as claude-haiku-5-5 and through Amazon Web Services, Google Cloud and Microsoft Azure. Claude Max and Team subscribers start receiving monthly Claude Platform API credits this week, with $100 on Max 5x, double that on Max 20x, and up to $500 shared on Team accounts usable across any Claude model. Adoption of tiered Opus-Sonnet-Haikupipelines in customer support and coding will signal confirmation.
