Apr 23, 2026 · View original article

OpenAI's GPT-5.5 targets long-horizon agents and knowledge work (April 2026)

Released 23 April 2026, GPT-5.5 is OpenAI's first new base model in a year, with a 400K context in Codex, strong agentic benchmarks, a high cyber and bio risk rating and tighter controls on sensitive requests.

On 23 April 2026 OpenAI released GPT-5.5, seven weeks after GPT-5.4. Where the March update refined an existing base, GPT-5.5 is described as OpenAI's first new base model in roughly a year, and its marketing centres on autonomy: the model is meant to plan, use tools, check its own work and keep going through ambiguous, multi-step tasks without constant direction. It launched in ChatGPT and Codex for Plus, Pro, Business and Enterprise customers, with API availability to follow. A GPT-5.5 Pro tier accompanies it, and Codex supports a 400K-token context.

OpenAI's reported results emphasise agentic and professional benchmarks. GPT-5.5 scored 82.7% on Terminal-Bench 2.0 for command-line coding, 73.1% on Expert-SWE for long-horizon software tasks, 84.9% on GDPval across 44 occupations, 78.7% on OSWorld-Verified for operating a computer, and 81.8% on CyberGym for security tasks. On GeneBench, a genetic data-analysis test, it reached 25.0%, a modest absolute figure the company presents as a sign of growing scientific utility. API pricing is set at $5 per million input tokens and $30 per million output tokens, double the GPT-5.4 rate, with GPT-5.5 Pro at $30/$180 and the usual 50% discount for batch and flex traffic.

The safety disclosures are notable. OpenAI rated the model high risk for both cybersecurity and biological capabilities under its preparedness framework, the first time both categories have been flagged at that level for a flagship release. Mitigations include stricter blocking of sensitive cyber requests, a Trusted Access programme that grants verified defensive-security users fuller capability, and pre-release external red-teaming and safety evaluations.

The release reflects a race that has compressed to weeks. Anthropic's Claude Mythos Preview, gated behind Project Glasswing on 7 April, and Claude Opus 4.7 on 16 April, set the bar OpenAI is answering, and Google's Gemma 4 and Meta's Muse Spark also arrived in the same month. TechCrunch framed GPT-5.5 as a step toward an OpenAI "super app" that combines chat, coding, computer use and documents in one product. For enterprise buyers, the substance is a model that is more capable of acting inside their systems, priced accordingly, and openly labelled as dual-use.

What it means for leaders

  • Match autonomy to controls. Long-horizon agents that self-verify still need scoped permissions, action logging and rollback; decide in advance which tasks may run unattended.
  • Re-run evaluations before switching. Benchmarks are OpenAI's own; test GPT-5.5 against your regulated use cases and document results as part of ISO/IEC 42001 change management.
  • Price the upgrade. Doubling input and output rates versus GPT-5.4 changes the economics of high-volume pipelines; consider routing simpler work to cheaper models.
  • Handle the dual-use rating formally. A vendor's high cyber and bio classification belongs in your risk register and in EU AI Act GPAI supplier due diligence, and may affect who is permitted to use the model internally.
  • Prepare for Trusted Access. Security teams that want full cyber capability for defence will need to enrol; align that with identity, monitoring and acceptable-use policies.

Comments

No comments yet. Be the first to comment.