
Anthropic Launches Claude Haiku 5.5 With 75% Cost Reduction
Anthropic officially launched its newest lightweight model, Claude Haiku 5.5, on Wednesday, October 7, 2026, marking the third major release within the Claude 5.5 family over the past month following the rollout of Opus and Sonnet.
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.
On average, it costs around 75% less to run than Claude Haiku 4.5. pic.twitter.com/jm06cZJfkV
– Claude (@claudeai) October 7, 2026
This strategic deployment significantly expands the company’s enterprise AI portfolio ahead of its planned Initial Public Offering (IPO). Designed to execute high-volume, cost-sensitive workloads and power autonomous agent workflows, the new model reduces average operational expenses by approximately 75% compared to its predecessor, Haiku 4.5.
New Pricing Cuts the Cost of Running the Model
Anthropic introduced a tiered pricing structure that sharply lowers the barrier to entry for enterprise scale. For prompts under 100,000 tokens, Haiku 5.5 is priced at $0.10 per million input tokens and $0.50 per million output tokens. This represents a substantial 90% reduction on input costs and a 50% cut on output rates compared to Haiku 4.5, which charged a flat rate of $1.00 per million input tokens and $5.00 per million output tokens regardless of request size.
For longer queries exceeding the 100,000-token threshold, the pricing adjusts to $0.50 per million input tokens and $2.50 per million output tokens. Anthropic noted that roughly 90% of developer requests to Haiku 4.5 naturally fell into the lower-priced tier, reflecting the high-frequency nature of tasks delegated to smaller models.
The revised fee structure includes:
- Prompts under 100,000 tokens: $0.10 per million input tokens / $0.50 per million output tokens.
- Prompts over 100,000 tokens: $0.50 per million input tokens / $2.50 per million output tokens.
- Prompt Cache Reads: $0.01 per million tokens for standard requests / $0.05 for extended requests.
- Prompt Cache Writes: $0.125 per million tokens for standard requests / $0.625 for extended requests.
Anthropic emphasized that the estimated 75% net workload savings accounts for both the price adjustments and an updated tokenizer that utilizes slightly more tokens per task to achieve higher fidelity.
What Tasks Was Haiku 5.5 Built to Handle?
Anthropic architected Haiku 5.5 to deliver sub-second responsiveness without compromising execution accuracy. The model reliably handles a broad spectrum of high-throughput operational tasks, from document summarization, data classification, and structured information extraction to database querying, content compaction, live customer support, voice agents, and autonomous browser navigation.
In addition to standalone operations, Haiku 5.5 is engineered to serve as an agile subagent alongside flagship models such as Sonnet 5.5 and Opus 5.5. In complex multi-agent architectures, larger models orchestrate high-level strategy while Haiku executes repetitive subtasks, preventing enterprise budgets from ballooning under full-tier model usage.
Financial AI company Rogo demonstrated this hybrid deployment in production. Alex Wang, who leads applied AI at Rogo, highlighted that while larger frontier models assemble end-to-end financial decks, Haiku 5.5 is tasked with retrieving discrete revenue figures for individual slides. Wang noted that the model delivers the necessary precision, speed, and cost efficiency required for continuous, high-density execution.
Furthermore, Haiku 5.5 introduces granular effort controls-a first for the Haiku tier. This feature allows developers to dial the model’s computational budget up or down depending on task complexity, with “medium” configured as the default setting.
Haiku 5.5 Delivers a Major Leap in Benchmark Performance
On technical benchmarks, Claude Haiku 5.5 scored 43 points on the Artificial Analysis Intelligence Index, placing it among the most capable lightweight models on the market. Anthropic’s internal evaluations demonstrate a decisive leap over Haiku 4.5 and competitive parity against flagship-adjacent models like OpenAI’s GPT-6 Luna, particularly across coding, computer control, and agentic workflows.

Evaluation results published by Anthropic highlight strong gains across key technical domains:
- OSWorld 2.1 (Computer Operation): Scored 72.4% on the offline subset, surging from 15.7% on Haiku 4.5 and outperforming GPT-6 Luna’s 48.9%.
- GDPval-AA v2.1 (Knowledge Work): Achieved 1,620 points, compared to 735 for Haiku 4.5 and 1,437 for GPT-6 Luna.
- AA-Briefcase v1.1 (Business Tasks): Reached 1,578 points, up from 614 on the previous generation.
- Terminal-Bench 4.0 (Command-Line Coding): Reached 39.2% at maximum effort, compared to 0.0% for Haiku 4.5 and 16.4% for GPT-6 Luna (scoring roughly 20% at default medium effort).
- Chartography (Visual Reasoning): Scored 46.4% without external tooling, up from 6.4% on Haiku 4.5.
- Humanity’s Last Exam (Multidisciplinary Reasoning): Attained 45.9% without tools and 57.4% when paired with tool use.
Does the Price Make Haiku 5.5 a Formidable Competitor?
Haiku 5.5 does not compete on raw reasoning alone; it directly targets the economics of enterprise AI deployment. At $0.10 per million input tokens and $0.50 per million output tokens, Haiku 5.5 matches OpenAI’s standard rates for GPT-6 Luna. However, an architectural distinction lies in the token threshold: Haiku’s higher tier kicks in past 100,000 tokens, whereas Luna maintains base rates up to 272,000 input tokens.
Across the broader landscape, Google offers Gemini 3.5 Flash-Lite at $0.30 for input and $2.50 for output per million tokens, alongside Gemini 3.8 Flash at promotional rates of $0.75 and $3.75. Meanwhile, xAI prices Grok 4.3 at $1.25 for input and $2.50 for output, scaling to $2.00 and $6.00 for Grok 4.7.
The pricing battle intensifies when compared against open and regional alternatives from Asia. Z.ai’s GLM-5.3-Flash closely mirrors Haiku 5.5’s reasoning benchmarks, slightly edging it out on GDPval-AA with 1,647 points versus 1,620 points. On the extreme low-cost end, Alibaba’s Qwen3.7 Flash starts at an ultra-low $0.03 input and $0.13 output per million tokens for prompts under 32,000 tokens.
While Haiku 5.5 is not the absolute cheapest model available globally, it establishes a compelling sweet spot combining aggressive cost reduction with top-tier agentic, command-line, and computer-use capabilities. For enterprises running mission-critical automation at scale, this balance delivers high throughput and reliability while keeping inference budgets firmly under control.
Enterprise Feedback and Real-World Velocity in Production
Early enterprise adoption reports indicate significant performance gains and reduced system latency in production environments. Yashodha Bhavnani, Vice President of AI Products at Box, reported an 11-point improvement in task evaluations over Haiku 4.5, alongside an approximate 50% reduction in response latency.
Similarly, Aaron Vinh, Staff Software Engineer at Asana, noted that end-to-end task completion latency dropped by more than 30%, with inference speed accelerating by up to 2.5 times per agent interaction. At HubSpot, Distinguished Software Engineer Ze’ev Klapow recorded a 92.8% success average across multiple evaluations within their CRM workflows, citing Haiku 5.5 as the strongest performer tested in its class.
These enterprise metrics affirm that Anthropic’s aggressive price cuts have not compromised model stability, ensuring low-latency data pipelines in demanding live applications.
Anthropic Slashes Sonnet Pricing and Adds Free API Credits
Alongside the Haiku release, Anthropic announced broader platform incentives to strengthen its developer ecosystem. The company cut Claude Sonnet 5.5’s cache-read pricing in half-from $0.20 down to $0.10 per million tokens. Anthropic estimates this change will reduce the running cost of typical agentic workflows by approximately 20% across supported cloud infrastructures.
To further encourage application development, Anthropic introduced recurring monthly API credits for premium account tiers on the Claude Platform:
- Max 5x Subscribers: $100 in monthly API credits.
- Max 20x Subscribers: $200 in monthly API credits.
- Team Subscribers: Up to $500 in pooled monthly API credits.
These credits apply across all models within the Claude family. Notably absent from the rollout was Fable 5.5 (belonging to Anthropic’s top-tier Mythos series), which remains subject to extended safety evaluations before enterprise release.
Haiku 5.5 Arrives with Tighter Security Controls and Wider Availability
Anthropic integrated upgraded cybersecurity safeguards into Haiku 5.5, outstripping the defensive baseline of Haiku 4.5. Standard configurations now automatically restrict malicious exploitation vectors, such as automated penetration testing. Organizations requiring expanded access for defensive security research or biological threat modeling can apply through Anthropic’s formal verification programs.
For immediate enterprise deployment, Haiku 5.5 is available via the Claude Platform under the model ID claude-haiku-5-5, as well as through major cloud hyperscalers including Amazon Web Services (Bedrock), Google Cloud (Vertex AI), and Microsoft Azure.
Complementing the launch, Anthropic updated its official Python and TypeScript Software Development Kits (SDKs) with beta support for autonomous computer interaction and web browsing, granting engineers the tooling needed to build responsive desktop and web agents.
The AI Race Shifts from Raw Power to Efficiency
The release of Claude Haiku 5.5 reflects a decisive strategic pivot across the artificial intelligence sector, moving beyond raw parameter scaling toward practical enterprise unit economics. With Anthropic positioning itself for a landmark public market debut, establishing a fast, reliable, and cost-effective model family is crucial for cementing long-term recurring revenue.
As enterprise IT leaders face mounting pressure to balance cutting-edge capability with strict cost controls, Haiku 5.5 offers a production-ready solution that blends speed and affordability, setting the stage for wider adoption of autonomous agent ecosystems across global industries.




