Anthropic has introduced Claude Haiku 5.5, a small AI model designed to deliver fast responses and lower operating costs for developers and businesses running AI-powered applications at scale. Announced on October 7, 2026, the model expands Anthropic’s Claude lineup with a more affordable option for high-volume workloads.
The biggest change is pricing. Anthropic estimates that Claude Haiku 5.5 costs around 75% less to run on average than Claude Haiku 4.5, while improving performance across several evaluated tasks. The company is targeting applications such as summarization, information extraction, classification, database queries, customer support, and browser-based automation.
However, the savings depend on the workload. Although Haiku 5.5’s standard input and output token prices are 90% lower than Haiku 4.5’s for prompts up to 100,000 tokens, the newer model’s tokenizer produces more tokens from the same text. That difference helps explain why Anthropic’s estimated average running-cost reduction is smaller than the headline price cuts.
Claude Haiku 5.5 focuses on speed and affordability
Anthropic describes Claude Haiku 5.5 as its fastest model at standard speed. It is designed for workloads where response times and operating costs matter, including customer support assistants, voice agents, document processing, and AI-powered software tools.
Haiku 5.5 can also handle narrower tasks within larger AI workflows. For example, a more capable model can manage a complex programming assignment while Haiku 5.5 extracts information from files, summarizes code, or retrieves specific details. Dividing work between models can help developers control the cost of running multi-step AI applications.
The model also supports adjustable reasoning effort, allowing developers to control how much effort it applies to a task. This provides another way to balance response speed, cost, and the complexity of the work.
Anthropic’s speed claims should be understood in context. Haiku 5.5 is positioned for fast inference at standard speed, but performance can vary with the task, workload, and operating mode. Anthropic notes that its Opus models can run faster in Fast Mode.
Claude Haiku 5.5 API pricing
Claude Haiku 5.5’s pricing is one of the most significant changes in this release. For prompts containing up to 100,000 tokens, Anthropic charges $0.10 per million input tokens and $0.50 per million output tokens.
| API pricing per 1 million tokens | Claude Haiku 5.5 | Claude Haiku 4.5 |
|---|---|---|
| Input tokens, prompts up to 100K | $0.10 | $1.00 |
| Output tokens, prompts up to 100K | $0.50 | $5.00 |
| Input tokens, prompts over 100K | $0.50 | — |
| Output tokens, prompts over 100K | $2.50 | — |
| Cache reads, prompts up to 100K | $0.01 | $0.10 |
| Cache reads, prompts over 100K | $0.05 | — |
For prompts up to 100,000 tokens, Haiku 5.5’s standard input and output rates are each 90% lower than Haiku 4.5’s. Cache reads are also 90% cheaper at this prompt size. For prompts exceeding 100,000 tokens, Haiku 5.5 has separate rates of $0.50 per million input tokens, $2.50 per million output tokens, and $0.05 per million cache-read tokens.
There is an important distinction between these per-token price reductions and the cost of processing the same material. Anthropic says its newer tokenizer produces approximately 30% more tokens for the same text than Haiku 4.5, although the difference varies by content.
As a result, a 90% reduction in token prices does not necessarily translate into a 90% reduction in the cost of completing a task. Anthropic’s estimate of around 75% lower average running costs accounts for differences in token usage and other workload considerations.
Anthropic is also reducing cached input-token prices for Claude Sonnet 5.5 by 50%, from $0.20 to $0.10 per million tokens. The company estimates that this change will make Sonnet 5.5 around 20% cheaper for most agentic workloads.
The company is also introducing monthly API credits for eligible Claude Max and Team subscribers. Max subscribers on the 5x plan receive $100 in monthly credits, while those on the 20x plan receive $200. Team subscribers receive up to $500 pooled across their users. These credits are intended to help subscribers build and experiment with applications and AI agents on Anthropic’s platform.
Claude Haiku 5.5 improves benchmark performance
Anthropic’s published evaluations show improvements over Haiku 4.5 across several areas, including computer use, knowledge work, reasoning, and coding.
On the OSWorld 2.1 computer-use benchmark, Claude Haiku 5.5 scored 72.4% on the offline subset, compared with 15.7% for Haiku 4.5. On Terminal-Bench 4.0, which evaluates agentic coding performance, Haiku 5.5 scored 39.2%, while Haiku 4.5 scored 0%.
These results come from Anthropic’s own evaluations and should be interpreted in that context. Benchmark scores can help compare models under defined testing conditions, but they do not guarantee equivalent performance across every real-world application.
Anthropic also shared early customer results. Asana reported a latency reduction of more than 30% for task completion in its AI Teammates product, with inference running up to 2.5 times faster per agent turn. HubSpot reported an average score of 92.8% across three runs on its CRM evaluation suite.
These figures reflect results reported by the respective companies in their testing environments, rather than independent measurements of performance across all workloads.
Haiku 5.5 is not intended to replace Anthropic’s larger models in every situation. Sonnet 5.5 and Opus 5.5 remain better suited to more complex agentic coding tasks, according to Anthropic. Haiku’s role is to handle smaller, well-defined jobs quickly and economically, particularly when a workflow generates many requests.
Claude Haiku 5.5 safety and availability
Anthropic says Haiku 5.5 improves on almost all of its alignment evaluations compared with Haiku 4.5, including evaluations related to misaligned behavior and cooperation with misuse.
The model also has updated safeguards for cybersecurity-related requests. Anthropic says its safeguards allow a wider range of defensive security work than those applied to Sonnet 5.5, while restricting activities it considers more likely to enable attackers. Its biological safety restrictions are aligned with those applied to several of the company’s other recent models.
The Claude Haiku 5.5 system card provides further information about the model’s safety evaluations and deployment safeguards.
Claude Haiku 5.5 is available through the Claude Platform and API, as well as Amazon Web Services, Google Cloud, and Microsoft Azure. Developers can access the model using the ID claude-haiku-5-5.
The model supports a one-million-token context window and up to 128,000 output tokens. This allows developers to process large inputs and generate lengthy responses when their applications require them.
With Haiku 5.5, Anthropic is making a stronger case for using smaller models in high-volume AI workflows. Lower token prices and improved benchmark results could make routine automation more affordable, although actual savings will depend on token usage, the task being performed, and performance in production.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
