A New Ceiling for AI Reasoning
OpenAI published benchmark results for o4-mini showing a 96.3% pass rate on AIME 2026, the competition mathematics dataset widely used to track reasoning-model progress. The prior record held by o3-mini was 85.1% on the same dataset, making the jump one of the largest single-generation gains since the reasoning model category was established.
The improvement stems from a combination of extended chain-of-thought training on verified proofs, a restructured reward model that penalises intermediate reasoning errors rather than only scoring final answers, and a 30% increase in total compute budget per query compared to o3-mini.
Performance vs. Cost Profile
Despite the capability jump, OpenAI priced o4-mini at the same API rate as its predecessor — $0.30 per million input tokens and $1.20 per million output tokens. The company frames this as a deliberate strategy to accelerate adoption in coding, scientific research, and financial modelling workflows where reasoning quality directly affects business outcomes.
On the company's internal CriticBench evaluation, o4-mini also outperformed GPT-5 on tasks requiring multi-step logical deduction, though GPT-5 retains advantages on open-ended instruction following and creative generation.
Practical Use Cases in Focus
OpenAI highlighted early access partners in pharmaceutical drug interaction screening, structural engineering verification, and derivatives pricing — domains where the 11-point accuracy gain on hard mathematical tasks translates directly to fewer costly errors. o4-mini is available in the API today; ChatGPT Plus integration follows in August.