← Defici Newsai-news

OpenAI o4-mini Posts Near-Perfect Score on AIME 2026, Widening Lead on Competition Math Benchmarks

By Defici Editorial · 26 Jul 2026

A New Ceiling for AI Reasoning

OpenAI published benchmark results for o4-mini showing a 96.3% pass rate on AIME 2026, the competition mathematics dataset widely used to track reasoning-model progress. The prior record held by o3-mini was 85.1% on the same dataset, making the jump one of the largest single-generation gains since the reasoning model category was established.

The improvement stems from a combination of extended chain-of-thought training on verified proofs, a restructured reward model that penalises intermediate reasoning errors rather than only scoring final answers, and a 30% increase in total compute budget per query compared to o3-mini.

Performance vs. Cost Profile

Despite the capability jump, OpenAI priced o4-mini at the same API rate as its predecessor — $0.30 per million input tokens and $1.20 per million output tokens. The company frames this as a deliberate strategy to accelerate adoption in coding, scientific research, and financial modelling workflows where reasoning quality directly affects business outcomes.

On the company's internal CriticBench evaluation, o4-mini also outperformed GPT-5 on tasks requiring multi-step logical deduction, though GPT-5 retains advantages on open-ended instruction following and creative generation.

Practical Use Cases in Focus

OpenAI highlighted early access partners in pharmaceutical drug interaction screening, structural engineering verification, and derivatives pricing — domains where the 11-point accuracy gain on hard mathematical tasks translates directly to fewer costly errors. o4-mini is available in the API today; ChatGPT Plus integration follows in August.

ShareXWhatsAppLinkedIn

Get Defici News in your inbox