OpenAI "o3-pro" API Goes Public — High-Reasoning Capabilities Unleashed for External Systems, Reshaping Legal and Research Workflows
機械翻訳 / Machine-translated
At 11:00 PM Japan time on August 31, 2026, OpenAI officially released the REST API for "o3-pro" — its reasoning-specialized model — to general developers. With capabilities previously exclusive to ChatGPT Pro subscribers ($200/month) now opened to the developer ecosystem, the pace at which AI penetrates "core professional work" — including automated review of lengthy legal documents, hypothesis generation for scientific papers, and complex financial modeling — is expected to accelerate significantly.
According to updates to OpenAI's official documentation, the endpoint is available under the model name o3-pro-2026-08-31. The context window is set at 200,000 tokens, with a maximum output of 32,768 tokens.
Pricing is $15.00 per 1M input tokens and $60.00 per 1M output tokens — 7.5 times that of the standard o3 (input $2.00 / output $8.00). However, with a BIG-Bench Hard (BBH) score of 97.3% and GPQA Diamond of 89.1%, it maintains an irreplaceable position in terms of accuracy.
"o3-pro has come to the API. It feels the same as when GPT-4 Turbo launched. We need to redesign not what we couldn't use before, but what we can now make possible going forward."
Access is restricted to accounts at "Tier 4" or above (cumulative usage of $250 or more), with a phased rollout for new developers.
o3 debuted as a research preview in December 2025 and became available via API in April 2026. However, the "pro" variant had long been kept outside the API as a differentiating feature exclusive to ChatGPT Pro.
This release is positioned as the third installment in the "agentic infrastructure build-out," following Anthropic's Claude Agent SDK and OpenAI's Operator API. Rather than a single model improvement, the trend of embedding high-reasoning capabilities into external workflows is accelerating.
Compared with Google Gemini 2.5 Ultra (BIG-Bench Hard 96.1%) and Anthropic Claude 4 Sonnet, o3-pro ranks higher in absolute accuracy for coding and mathematics, though Claude 4 Sonnet is considered superior in cost efficiency.
An entire legal document (typically 150,000–300,000 characters) can be processed in one pass. While RAG or chunked processing was previously mandatory, architectures that feed an entire contract plus a case law database in a single request become realistic.
At $15/M input tokens, this is a level reserved for "high-value judgment tasks only." The price point reads as targeting a market for professional substitution — distinct from coding copilots or summarization tools. The industry has entered a phase of weighing cost against operational impact.
Eligibility requires reaching Tier 4 (cumulative usage of $250 or more), which means new startups face a warm-up period of several weeks. It is worth noting that the structure gives established heavy users a first-mover advantage.
The 97.3% figure published by OpenAI is based solely on their own measurements. Independent benchmark organizations (HELM, LMSYS Arena, etc.) have yet to release verification results, meaning continued reporting is needed before making production deployment decisions.
Looking only at the price may lead to the conclusion that it offers poor value — but the real point is that the ability to identify which tasks to apply it to becomes a competitive advantage.
Comparing the hourly rate of a single attorney against o3-pro API costs, there are clearly domains where even a high price point is economically justified. Review of pharmaceutical regulatory documents, evidentiary review in accounting audits, and due diligence for investment contracts all fall within the same logic.
On the other hand, thoughtless integration — "let's just use o3-pro" — carries the risk of monthly API costs immediately reaching six figures (in yen). This is also a moment when engineers who can design model selection at the architecture level become increasingly scarce and valuable.
The "brake on adoption" built into the Tier restrictions is also worth noting. OpenAI's strategy of managing infrastructure load while releasing capabilities in stages is clearly visible.
The structural significance of "the highest-reasoning model entering external systems" has not yet been fully priced in by the market as a whole. The next focal points are independent benchmark verification results and the timing of Tier restriction easing. Organizations that can design which tasks to apply it to are expected to be the first to generate a meaningful revenue gap.
This article was written by an AI writer (AI News) from the Mirai News editorial team.