Claude 5 Haiku Announced — Claude 4 Sonnet-Level Reasoning at 75% Less
機械翻訳 / Machine-translated
機械翻訳 / Machine-translated
On September 2, 2026, just past 11 PM Japan time, Anthropic quietly released "Claude 5 Haiku" without any prior announcement. Priced at $0.75 per million input tokens and $3.75 per million output tokens — a 75% price reduction compared to Claude 4 Sonnet — the model nonetheless outperforms its predecessor on coding and mathematical reasoning benchmarks. This marks the day the design assumption that "cheaper equals lower performance" fell apart.
Anthropic simultaneously updated its official blog and API documentation, making the model available immediately. Key benchmark results are as follows:
The context window remains at 2 million tokens. Average response speed has improved by approximately 35% compared to Claude 4 Haiku.
"Claude 5 Haiku has made the dream of Sonnet-quality at bargain pricing a reality. This will change how production teams select APIs." (AI engineer, via X)
In the Claude 4 series, the performance gap between Haiku and Sonnet was clear-cut — Sonnet or above was effectively required for complex reasoning tasks. While Haiku excelled in speed and cost, it fell short of Sonnet in coding accuracy, limiting its use in production environments.
In the first half of 2026, inference costs fell sharply across the industry — estimated at an average drop of 60–70%. As companies competed to bring high performance down to lower-priced model tiers, Anthropic has now delivered a clear answer to that challenge.
On the competitive front, recent performance gains from DeepSeek R3 and Gemini 2.5 Ultra have been notable, and refreshing its lightweight model lineup had reportedly become a pressing priority for Anthropic as well.
For SaaS products or chatbots handling around 100 million requests per month, switching from Claude 4 Sonnet to Claude 5 Haiku could reduce monthly API costs by up to three-quarters. Equivalent performance at one-quarter the cost is a figure that directly affects a product's revenue model.
The architectural shift from "batch processing with high-cost models" to "massively parallel execution with lightweight, high-performance models" is expected to accelerate. The scenario of Claude 5 Haiku running in parallel as a subagent has now become a viable option from a cost perspective.
This release covers Haiku only — announcements for Claude 5 Sonnet and Opus have yet to come. Anthropic's strategy of rolling out models from the lower tier upward is now unmistakable. When Sonnet arrives, its performance and pricing bar will need to be read in light of these results.
What stands out is the "no-announcement, same-day release" approach. With no prior leaks or teasers, API documentation and the blog are updated simultaneously. This stands in contrast to the drip-feed approach of OpenAI and Google — Anthropic consistently follows a style of "ship first, talk later." It appears to be a deliberate choice to let developers experience the product directly, rather than maximize announcement impact.
On the performance side, the fact that Claude 5 Haiku surpassed Claude 4 Sonnet on HumanEval should not be overlooked. A lightweight model beating a mid-tier model on coding has direct implications for how tools like GitHub Copilot and similar products select their backend models.
In the Japanese market, AI-powered customer service and document processing services built on the Claude API are growing rapidly. The impact on these cost structures is expected to be felt at the operational level within this week.
Claude 5 Haiku has refuted with hard numbers the assumption that "lightweight models are weak at reasoning." The democratization of inference costs advances further, and we are now entering a phase where three core assumptions — agent design, product economics, and model selection — are being called into question simultaneously.
The next focal point is the timing of Claude 5 Sonnet's release. If Haiku sets this bar, how high will Sonnet reach? The LLM price-versus-performance competition of autumn 2026 is still only in its opening chapter.
This article was written by an AI writer (AI News) from the Mirai News editorial team.