Anthropic releases Claude Haiku 5.5 with 1M token context and adaptive thinking
In short: Anthropic has launched Claude Haiku 5.5 (claude-haiku-5-5), positioned as its most capable Haiku model for high-volume, latency-sensitive work. It supports a 1 million token context window, up to 128k max output tokens, and adaptive thinking controlled via a new effort parameter. The model is available on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry. Anthropic warns that code written for the previous Haiku 4.5 model may break on the new version.
This summary was generated automatically by AI from Claude Platform Release Notes's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.
What changed?
- 1New model claude-haiku-5-5 released, tuned for high-volume and latency-sensitive tasks
- 21M token context window and 128k max output tokens
- 3Adaptive thinking enabled via a new effort parameter, replacing manual budget_tokens control
- 4Available across Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry
- 5Breaking change: manual extended thinking (budget_tokens) now returns a 400 error, adaptive thinking is on by default, responses may start with thinking blocks, and the same text can consume more tokens
| Parameter | Before | Now |
|---|---|---|
| Context window | Not specified (Claude Haiku 4.5) | 1,000,000 tokens |
| Max output tokens | Not specified | 128,000 tokens |
| Reasoning / thinking control | Manual extended thinking via budget_tokens parameter | Adaptive thinking on by default, controlled via effort parameter |
| Availability | Claude Haiku 4.5 on Claude API, Bedrock, AWS, Google Cloud, Microsoft Foundry | Claude Haiku 5.5 on Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, Microsoft Foundry |
Why it matters
A large context window combined with a model built for speed and volume is directly relevant to teams processing long conversations or documents under tight latency budgets, such as real-time agents or high-throughput API pipelines. The breaking changes mean anyone upgrading from Haiku 4.5 needs to review their thinking-related code and token budgeting before switching models.
What it means for AI agents and contact centers
For teams running voice agents or tool-calling pipelines on Claude, this model is worth evaluating for latency and cost on high-volume call flows, but existing integrations built on Haiku 4.5's manual extended thinking (budget_tokens) will need code changes before migrating, since that parameter now errors out and adaptive thinking runs by default.
🧪 Worth Testing
Its combination of a 1M token context window, adaptive thinking via the effort parameter, and tuning for latency-sensitive, high-volume work makes it a strong candidate to benchmark against current models for voice agent and tool-calling use cases, including cost and response speed.
Sources
- Claude Platform Release NotesOfficialPrimary sourceOriginal article →„We've launched Claude Haiku 5.5 ( claude-haiku-5-5 ), our most capable model tuned for high-volume and latency-sensitive work. It has a 1M token context window , 128k max output tokens, and adaptive t“7 Oct 2026, 03:00
- AWS Machine LearningOfficialOriginal article →„Introducing Claude Haiku 5.5 on AWS“7 Oct 2026, 21:52
- Published by source
- 7 Oct 2026, 03:00
- Found by our system
- 7 Oct 2026, 21:06
- Summary generated
- 7 Oct 2026, 21:08
This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy