OpenAI Debuts GPT-5.4 Pro and Thinking Models with Million-Context Window
elementOpenAI has officially announced the release of its latest foundational model, GPT-5.4 , which it describes as the most capable and efficient professional-grade model to date. According to AIbase, the series follows a differentiated launch strategy: alongside the standard version, OpenAI introduced GPT-5.4Thinking—a reasoning model specialized in complex logic—and GPT-5.4Pro, built for high-performance tasks.

On the technology front, the API version of GPT-5.4 delivers a major upgrade, featuring a context window of up to 1 million tokens—the largest ever offered by OpenAI. The model also achieves notable gains in token efficiency, enabling it to solve similar problems with fewer resources.
In safety and accuracy, the new model reduces the per-statement error rate by 33% compared to GPT-5.2, and cuts overall response errors by 18%. To mitigate potential "chain-of-thought deception" risks in reasoning models, OpenAI has introduced a new security evaluation system. Tests indicate that GPT-5.4Thinking offers greater transparency, making it difficult to conceal or fabricate its reasoning steps.
In benchmark evaluations, GPT-5.4 delivered strong results, setting new records in computer usage tests like OSWorld-Verified and WebArena Verified, while also achieving an impressive 83% on the GDPval knowledge task.
Mercor CEO Brendan Foody noted that the model also leads the APEX-Agents benchmarks in professional domains like finance and law, particularly excelling at generating financial models, legal analysis, and other long-form deliverables. With the new "tool search" system, the model becomes more efficient when invoking external tools, dramatically reducing token overhead in large-scale tool integration scenarios.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500

On the technology front, the API version of
In safety and accuracy, the new model reduces the per-statement error rate by 33% compared to GPT-5.2, and cuts overall response errors by 18%. To mitigate potential "chain-of-thought deception" risks in reasoning models,
In benchmark evaluations,
Mercor CEO Brendan Foody noted that the model also leads the
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage





Home






