Claude 5 Sonnet Launch — Entering the Agent Cost Race at $3/MTok with 200K Context
機械翻訳 / Machine-translated
機械翻訳 / Machine-translated
On September 4, 2026, Anthropic officially released "Claude 5 Sonnet." Positioned as a "core model" that maintains the low-price strategy established with Claude 5 Haiku while significantly improving accuracy and context length, the API was opened to all tiers on launch day — a direct challenge to the enterprise market where o3-pro and Gemini 2.5 Ultra are locked in fierce competition.
The specifications published in Anthropic's official blog and API documentation are as follows:
Immediately after the announcement, reactions from engineers flooded X.
HumanEval 94.7%. That puts it on par with o3-pro for code generation, and Anthropic wins hands down on price. Looks like our team will be debating tonight about where to put the orchestrator layer for our agents.
Rollout to Claude.ai Pro users began at 10:00 PM JST on the same day, and API access is available immediately starting from Tier 1.
Anthropic announced Claude 5 Haiku in June 2026, delivering Claude 4 Sonnet–level reasoning at a 75% cost reduction. If Haiku is the "practical workhorse for speed and low cost," this Sonnet is the "flagship model optimizing the triangle of accuracy, speed, and cost." The three-tier stack of Opus (highest accuracy) · Sonnet (balanced) · Haiku (fast and low-cost) has once again become clearly defined.
Claude 5 Opus has yet to receive an official announcement and is expected to launch before the end of the year. Anthropic's strategy is to control pressure through a phased rollout while positioning Sonnet as "the most cost-effective choice available right now."
Enterprise AI agents consume tens of thousands to hundreds of thousands of tokens per task. The $3/MTok input pricing is approximately 14% below Gemini 2.5 Pro's $3.5/MTok, and roughly 40% cheaper than GPT-4o–equivalent products. For large-scale agent operations where monthly API costs run into the millions of yen, this difference has a direct impact on ROI calculations.
The thinking token limit has been expanded fourfold from Haiku's 8,192 to 32,768. Accuracy improvements are expected for tasks requiring "deep deliberation," such as cross-referencing legal document provisions and sensitivity analysis in financial models. Anthropic has stated in internal evaluations that using this mode yields an additional 4–6 point improvement on challenging benchmarks.
Video input is supported via a frame-by-frame extraction and submission method. Integration into use cases such as manufacturing line anomaly detection, video content QA workflows, and surveillance camera analysis is now feasible at realistic costs.
Access across all tiers is available from day one, and rate limits have been relaxed to 1.5× those of Claude 3.7 Sonnet. The barriers to adoption for batch processing and high-frequency agent calls have been clearly lowered.
The most important aspect of this announcement is the "restructuring of the model stack." With a core model now placed between Haiku and Opus, developers can redesign their criteria for selecting models by use case. The optimal solution for agent design is likely to be a two-tier structure of "orchestrator = Sonnet, worker = Haiku." For companies seeking both cost efficiency and accuracy, this combination is expected to become the de facto standard configuration in the second half of 2026.
The intent behind the pricing is also readable. The $3/MTok figure is deliberately set just below Gemini 2.5 Pro's level. It is a strategy of "price optimization rather than price disruption" — aiming to be chosen in price comparisons while holding its own on accuracy as well.
From a Japanese enterprise perspective, the Enterprise plan's 1M token context expansion becomes a practical option. This is a moment where it is worth immediately revisiting cost estimates for batch processing workflows involving lengthy specifications, multiple meeting minutes, and regulatory documents.
Claude 5 Sonnet fills the "missing core" in Anthropic's model lineup and raises its competitiveness in the enterprise AI agent market by another level. The next focus is when Claude 5 Opus will arrive — and how far it will be able to break o3-pro's monopoly on high accuracy.
This article was written by an AI writer (AI News) from the Mirai News editorial team.