Dec 11, 2025 · View original article

OpenAI Releases GPT-5.2 With Three Variants and a 30% Cut in Hallucinations

Released on 11 December 2025 after an internal 'code red', GPT-5.2 arrives in Instant, Thinking and Pro variants with 256k context, a $1.75 per million input price and new leads on GDPval and ARC-AGI-2.

On 11 December 2025 OpenAI released GPT-5.2, roughly a month after GPT-5.1 and three weeks after Google's Gemini 3 and Anthropic's Claude Opus 4.5 had displaced OpenAI from the top of most public leaderboards. The launch followed reports that CEO Sam Altman had declared an internal "code red" to concentrate resources on the core model; the company positioned GPT-5.2 as its answer, and Altman told CNBC he expected OpenAI to be out of code red by January.

The model ships in three variants. GPT-5.2 Instant targets everyday tasks, GPT-5.2 Thinking handles complex reasoning and is the default for most professional work, and GPT-5.2 Pro provides maximum compute for the hardest problems. OpenAI's headline figures are 70.9% (Thinking) and 74.1% (Pro) on GDPval, its benchmark of economically valuable professional tasks judged against human experts; 55.6% on SWE-bench Pro and 80.0% on SWE-bench Verified; 92.4% to 93.2% on GPQA Diamond; 52.9% to 54.2% on ARC-AGI-2; 40.3% on FrontierMath tiers one to three; and a perfect score on AIME 2025. OpenAI also reports that GPT-5.2 Thinking produces about 30% fewer hallucinations than GPT-5.1 on its factuality evaluations and behaves better in conversations touching on mental health.

Commercially, the API price rises to $1.75 per million input tokens and $14 per million output tokens, with cached input at $0.175, and GPT-5.2 Pro at $21 and $168 respectively. The context window is 256k tokens, and a new /compact endpoint in the Responses API is designed for long agentic sessions. The model rolled out to paid ChatGPT plans and the API on launch day.

The context is unusual. For most of 2025 OpenAI set the pace; in November it found itself responding to competitors on their timeline. GPT-5.2's benchmark claims restore rough parity rather than a decisive lead: its SWE-bench Verified score edges Gemini 3 Pro's 76.2%, and its ARC-AGI-2 figure surpasses Gemini 3 Deep Think's 45.1%, but Anthropic's Opus 4.5 remains highly competitive on agentic coding. The more durable signal is the focus on GDPval and professional tasks, which reflects the industry's pivot from consumer chat toward enterprise work automation, where reliability, cost and auditability matter more than raw scores.

For buyers, the price increase is worth noting. After a year in which frontier prices mostly fell, OpenAI raised list prices for its flagship while arguing that improved token efficiency lowers total cost per task. Whether that holds will depend on workload, and it strengthens the case for measuring cost per completed task rather than cost per token.

What it means for leaders

  • Benchmark against your own tasks. GDPval-style professional evaluations are useful, but the numbers that matter are accuracy and cost on your documents, code and workflows; build a small internal evaluation set and rerun it on every model release.
  • Model cost per outcome. Higher token prices combined with lower token consumption require a total-cost analysis; procurement should request token-usage data from pilots before renewing volume commitments.
  • Hallucination reductions do not remove human oversight. A 30% improvement still leaves material error rates for high-stakes uses; keep review controls in place as NIST AI RMF and the EU AI Act's human-oversight expectations require.
  • Expect release cycles of weeks, not years. Governance processes that take a quarter to approve a model version will lag the market; define a fast-track re-assessment for point releases from already-approved vendors.
  • Log which variant is in use. Instant, Thinking and Pro have different capabilities and prices; record the variant in your AI system inventory and audit trails.

Comments

No comments yet. Be the first to comment.