When OpenAI shipped GPT-5.5 on April 23, 2026, the headlines said the same thing: new king of the benchmarks. Artificial Analysis crowned it the leading model, three points ahead of a three-way tie. But I read past the launch tables, and the real story is messier, more expensive, and already half out of date.
TL;DR: GPT-5.5 is OpenAI's flagship model, released April 23, 2026, at $5 input / $30 output per 1M tokens (double GPT-5.4). It topped the Artificial Analysis Intelligence Index at 60 on launch day and leads on agentic and shell-driven work like Terminal-Bench 2.0 (82.7%). But Claude Opus 4.7 still beats it on 6 of 10 shared benchmarks, hallucinates far less (36% vs 86% on AA-Omniscience), and costs $5 less per million output tokens. As of June 2026, the index crown has already moved to Opus 4.8 (Claude Fable 5 briefly topped it in early June, but Anthropic suspended it on June 12 after a US export-control directive). If you do long-running agent work, GPT-5.5 earns its price. For deep reasoning or factual accuracy, it does not.
Here's the thing about "smartest model yet" claims: they have a shelf life measured in weeks. So this is not a launch hype piece. I'm going to walk you through what GPT-5.5 actually costs, where it genuinely wins, where independent testers caught it falling short, and whether the price hike makes sense for the work you do.
If you want the family map first, our GPT-5.4 family explainer covers the model OpenAI just doubled the price over.
What GPT-5.5 actually is
GPT-5.5 is OpenAI's frontier model family, announced April 23, 2026. It ships in three flavors you'll actually meet: GPT-5.5 Thinking (the reasoning workhorse in ChatGPT and the API), GPT-5.5 Pro (a higher-effort variant for Pro and Enterprise), and GPT-5.5 Instant (the fast default that quietly replaced GPT-5.3 Instant in ChatGPT on May 5, 2026).
Thinking and Pro went live in ChatGPT and Codex on launch day. The API followed one day later, on April 24, with OpenAI saying it held it back 24 hours for "different safeguards."
GitHub Copilot got access the same day as the API. The internal codename, if you collect trivia, was "Spud."
One claim shows up in nearly every independent review: GPT-5.5 is the first ground-up retrained base model since GPT-4.5, with the entire 5.1 through 5.4 series being refinements on the same weights. Two reviewers (NewforTech and Awesome Agents) state this flatly.
I want to flag it carefully, though. OpenAI's own announcement never uses the word "retrain," and a six-week gap from GPT-5.4 would be unusually fast for a full retrain. So treat it as the prevailing read of the launch, not a confirmed engineering fact.
OpenAI's GPT-5.5 family (Thinking, Pro, and the GPT-5.5 Instant default) at $5/$30 per 1M tokens, with leads on agentic and shell-driven work
Best for: General knowledge workers needing a capable all-in-one assistant, Developers wanting quick code generation with GPT Image and Codex integration
Pro Tip: If you only use ChatGPT casually, you're probably already on GPT-5.5 Instant without knowing it. It became the default for all users on May 5, 2026. Paid users can keep selecting GPT-5.3 Instant from the model picker for 3 months before OpenAI retires it, so check your picker now if your prompts suddenly feel different.
GPT-5.5 pricing: the $5/$30 reality
Let's talk money, because this is where GPT-5.5 gets controversial. Standard GPT-5.5 lists at $5 input / $30 output per 1M tokens. That's a clean doubling of GPT-5.4's $2.50/$15.
GPT-5.5 Pro sits in a different bracket entirely at $30 input / $180 output. Batch and Flex modes run at half price; Priority runs at 2.5x.
A doubled price tag sounds brutal. But there's a wrinkle that softens it, and a counter-wrinkle that brings the pain right back.
Why the price hike is not quite 2x in practice
Artificial Analysis found GPT-5.5 (xhigh) uses roughly 40% fewer output tokens to run their Intelligence Index than GPT-5.4 did. Because output tokens are the expensive half, that token efficiency absorbs most of the price increase.
Net result: running the AA Index on GPT-5.5 costs about 20% more than GPT-5.4, not 100% more. OpenAI leans on this "fewer tokens" framing too, especially for Codex tasks.
The efficiency story doesn't always hold. Tessl ran 1,742 tests with engineering skills loaded into the prompt and measured GPT-5.5 at $0.49 per run versus GPT-5.4 at $0.30, a 63% premium. And the score? 89.4 versus 89.3. Within a rounding error. Once skill-prompts enter the picture, you can pay 63% more for the same answer.
How GPT-5.5 pricing compares to rivals
Here's where the pricing gets uncomfortable for OpenAI. Claude Opus 4.7 matches GPT-5.5 on input ($5) but charges $5 less on output ($25 vs $30), and it ships a 1M-token context window too.
So GPT-5.5 is the more expensive model on output while trailing Opus on most reasoning benchmarks. Gemini 3.1 Pro undercuts both at $2 input / $12 output.
Claude Opus 4.7 prices output $5 cheaper at $5/$25, hallucinates far less, and leads GPT-5.5 on 6 of 10 shared reasoning and code-review benchmarks
Best for: Knowledge workers doing deep document analysis and long-form writing, Developers who need precise instruction-following for complex multi-step tasks
| Model | Input / 1M | Output / 1M | Context | Released |
|---|---|---|---|---|
| GPT-5.5 | $5.00 | $30.00 | ~1M | Apr 23, 2026 |
| GPT-5.5 Pro | $30.00 | $180.00 | ~1M | Apr 23, 2026 |
| Claude Opus 4.7 | $5.00 | $25.00 | 1M | (prior) |
| Gemini 3.1 Pro | $2.00 | $12.00 | 1M+ | (prior) |
There's a smarter way to read this than headline rates, though. Artificial Analysis showed that GPT-5.5 on its "medium" effort setting hits the same AA Index score as Opus 4.7 on "max" effort, for about a quarter of the cost (roughly $1,200 versus $4,800 to run the full index). Useful.
But Gemini 3.1 Pro reaches that same parity for around $900. So even the price-favorable reading of GPT-5.5 doesn't make it the cost leader. Gemini still is.
Gemini 3.1 Pro undercuts both on price at $2/$12 and reaches frontier-parity on the AA Index for the lowest cost of the three
Best for: Google Workspace users wanting native AI inside Docs, Gmail, and Sheets, Researchers needing citation-backed outputs via NotebookLM
Pro Tip: Drop the reasoning effort to "medium" before you reach for a cheaper model. GPT-5.5 (medium) matched Opus 4.7 (max) on the AA Index at about 1/4 the cost in independent testing. You buy frontier-parity output without paying flagship effort prices, and you keep the OpenAI tooling you already wired up.
Performance: where GPT-5.5 genuinely wins
On launch day, Artificial Analysis put GPT-5.5 (xhigh) at the top of its Intelligence Index with a score of 60, breaking a three-way tie with Anthropic and Google. The median across 141 models they track is 33, so this is real frontier territory. Three points is a lead, not a blowout, but it was a lead.
The clearer story is where GPT-5.5 wins, and it's a specific shape. Its strongest result by far is Terminal-Bench 2.0 at 82.7%, a 13-point gap over Opus 4.7 (69.4%) and Gemini 3.1 Pro (68.5%). That benchmark tests real command-line workflows: planning, iteration, and coordinating tools across many steps.
GPT-5.5 also leads on BrowseComp (84.4%), OSWorld-Verified (78.7%, basically a tie), and CyberGym (81.8%).
See the pattern? GPT-5.5's wins cluster on long-running tool use, shell-driven automation, and computer-use tasks. The work where a model has to keep going for many steps without losing the thread.
Cursor's CEO Michael Truell put it well: "stays on task significantly longer without stopping early."
Long context and abstract reasoning gains
Two areas improved a lot over GPT-5.4. Long-context recall roughly doubled: MRCR v2 8-needle recall at 512K to 1M tokens jumped from 36.6% to 74.0%. And on ARC-AGI-2, the abstract-reasoning test, GPT-5.5 leads at 85.0% against Opus 4.7's 75.8% and Gemini's 77.1%.
But notice the asterisks. On ARC-AGI-1, Gemini still wins (98.0% vs 95.0%). And on Graphwalks parents at 1M tokens, Opus still beats GPT-5.5 (72.0% vs 58.5%). Even the wins come with neighbors that lose.
Pro Tip: If your workload is shell automation or multi-step agent loops, run a 20-task pilot on GPT-5.5 before committing. The Terminal-Bench 2.0 lead (82.7% vs 69.4%) is the one place the price premium most clearly pays for itself, and a small pilot tells you in an afternoon whether it holds for your specific tooling.
Where GPT-5.5 falls short
Now the part the launch posts skipped. Three independent findings should make you pause before you default to GPT-5.5 for serious knowledge work.
First, hallucinations. On Artificial Analysis's private AA-Omniscience benchmark, GPT-5.5 (xhigh) records the highest accuracy they've ever measured (57%), but at an 86% hallucination rate. Compare that to Opus 4.7 at 36% and Gemini 3.1 Pro at 50%.
In plain terms: when GPT-5.5 doesn't know an answer, it's far more likely to make one up than its rivals are. OpenAI's own announcement never addresses this number.
Second, the benchmark scoreboard is not the sweep the headlines implied. Across 10 shared OpenAI and Anthropic benchmarks, Opus 4.7 leads on 6: SWE-Bench Pro (64.3% vs 58.6%), HLE with and without tools, GPQA Diamond, MCP Atlas, and FinanceAgent v1.1.
GPT-5.5 leads on 4. So on reasoning-heavy and code-review-grade work, the model you actually want is often Opus.
Third, head-to-head reviews split hard depending on the rubric. Tom's Guide ran a 7-category comparison and GPT-5.5 lost all 7 to Opus 4.7, praising its speed but criticizing the hallucinations. ZDNET, on the other hand, gave it 93/100 across 10 rounds, docking points "only for exuberance."
Reviewer methodology swings the result, so take any single verdict with salt, including ours.
GPT-5.5: What Works and What Doesn't
What Works
- Best-in-test on agentic and shell work: Terminal-Bench 2.0 at 82.7%, a 13-point lead
- Long-context recall roughly doubled vs GPT-5.4 (74.0% MRCR at 512K-1M)
- ~1.5x faster end-to-end than GPT-5.4 in Tessl's independent test (89.5s vs 135.4s per run)
- Token efficiency softens the price hike to ~20% more, not 100%, on the AA Index
- Strong cyber-defense capability: UK AISI called it possibly the strongest model they have tested on expert tasks
What Doesn't
- 86% hallucination rate on AA-Omniscience, far above Opus 4.7 (36%) and Gemini 3.1 Pro (50%)
- Trails Opus 4.7 on 6 of 10 shared benchmarks, including code-review-grade SWE-Bench Pro by 5.7 points
- $30 output is $5 more than Opus 4.7 and $18 more than Gemini 3.1 Pro
- "First full retrain" claim is reviewer-reported, not confirmed by OpenAI
- AISI red-teamers found a universal jailbreak in 6 hours, and the final safeguard fix was never fully verified
GPT-5.5 Instant: the default-model swap you didn't choose
On May 5, 2026, OpenAI swapped ChatGPT's default model to GPT-5.5 Instant for everyone, skipping a "GPT-5.4 Instant" entirely. According to OpenAI's internal evaluations, Instant cuts hallucinated claims by 52.5% on high-stakes prompts (medicine, law, finance) and 37.3% on conversations users had flagged as wrong.
HealthBench rose from 49.6 to 51.4. AIME 2025 jumped from 65.4 to 81.2.
Behind the scenes, Instant auto-routes: for any given prompt, it decides whether to answer with the lighter GPT-5.3 Instant or escalate to GPT-5.5 Thinking. There's also a new "memory sources" panel that shows which past chats, files, or Gmail items shaped a personalized answer, with controls to delete or correct each one.
Two honest caveats here. Every Instant number above is OpenAI's own internal eval against its own predecessor: no independent benchmark, and no comparison to Google or Anthropic. And the memory-sources panel admits it "may not show every factor that shaped an answer," which a HiddenLayer security exec warned creates a second context log that can conflict with the audit trails enterprises already keep. Useful, but not the full picture.
The cyber capability nobody expected
This is the result that genuinely surprised me. The UK AI Security Institute tested GPT-5.5 on expert-level cyber tasks and scored it at 71.4%, ahead of Anthropic's restricted Mythos Preview (68.6%), GPT-5.4 (52.4%), and Opus 4.7 (48.6%). AISI wrote that "on this measure, GPT-5.5 may be the strongest model we have tested."
One spotlight makes it concrete. On a reverse-engineering challenge that took an expert human about 12 hours with professional tools, GPT-5.5 with a basic agent in a Kali Linux container solved it in 10 minutes 22 seconds for $1.73 in API usage, with no human help. It diagnosed relocation tables, built an emulator, caught its own swapped read/write bug, and constraint-solved the password.
That cuts both ways, and AISI said so. Their red-teamers found a universal jailbreak in 6 hours that worked across every malicious cyber query OpenAI gave them. OpenAI patched it, but AISI couldn't verify the final fix because of a config issue in the version they received.
So the capability is real, and so is the open safety question.
How GPT-5.5 stacks up right now
Here's the freshness check that the launch coverage can't give you. GPT-5.5 was the leading model on April 23. As of June 2026, it isn't anymore.
Anthropic shipped Claude Opus 4.8 on May 28, and it sits above GPT-5.5 on the Artificial Analysis index (56 vs 55). Anthropic did briefly ship Claude Fable 5, a Mythos-class model, in early June, but on June 12 it suspended access to Fable 5 and Mythos 5 for all customers worldwide after a US government export-control directive flagged a jailbreak concern. So Fable 5 is not something you can pick today. Opus 4.8 is, and it is the model that took the crown. The "smartest model yet" label lasted about five weeks.
GPT-5.5 is still OpenAI's current flagship, though. A GPT-5.6 has been rumored for a June 2026 release (prediction markets put the odds high), but as of this writing OpenAI has made no official announcement, so GPT-5.5 remains the model you're actually choosing today. For the wider picture, see our AI tools coverage, the model comparison hub, and how its rivals handle workspace agents.
| Use case | Best pick | Why |
|---|---|---|
| Shell automation / long agent loops | GPT-5.5 | 82.7% Terminal-Bench 2.0, 13-point lead, stays on task longer |
| Deep code review / reasoning | Claude Opus 4.7 | Leads 6 of 10 benchmarks, far lower hallucination rate |
| Factual knowledge work | Claude Opus 4.7 | 36% hallucination vs GPT-5.5's 86% on AA-Omniscience |
| Lowest cost at frontier parity | Gemini 3.1 Pro | ~$900 to run the AA Index vs GPT-5.5's ~$1,200 |
| Casual ChatGPT use | GPT-5.5 Instant | Already your default, shorter answers, fewer hallucinations than 5.3 |
Frequently asked questions
How much does GPT-5.5 cost?
GPT-5.5 costs $5 per 1M input tokens and $30 per 1M output tokens via the API, double GPT-5.4's $2.50/$15. GPT-5.5 Pro is far pricier at $30/$180. Batch and Flex modes run at half rate. In ChatGPT, GPT-5.5 Thinking is on Plus and above, and GPT-5.5 Instant is the free default.
Is GPT-5.5 better than Claude Opus 4.7?
It depends on the task. GPT-5.5 wins on agentic and shell-driven work like Terminal-Bench 2.0. Opus 4.7 leads on 6 of 10 shared benchmarks, hallucinates far less (36% vs 86% on AA-Omniscience), and costs $5 less per million output tokens. For deep reasoning and factual accuracy, Opus is the safer pick.
Is GPT-5.5 still the leading AI model?
No, not anymore. GPT-5.5 topped the Artificial Analysis Intelligence Index at 60 when it launched on April 23, 2026, but Claude Opus 4.8 (May 28) has since overtaken it (56 vs 55). Claude Fable 5 briefly led in early June, but Anthropic suspended it for all customers on June 12 after a US export-control directive, so it is not currently available. GPT-5.5 remains OpenAI's current flagship, with no GPT-5.6 officially announced as of June 2026.
What is GPT-5.5 Instant?
GPT-5.5 Instant is the fast default model that replaced GPT-5.3 Instant in ChatGPT on May 5, 2026. It auto-routes between a light model and GPT-5.5 Thinking per prompt, gives shorter answers, and (per OpenAI's internal evals) cuts hallucinations 52.5% on high-stakes prompts versus GPT-5.3 Instant.
Does GPT-5.5 hallucinate a lot?
On the AA-Omniscience benchmark, yes: GPT-5.5 (xhigh) records an 86% hallucination rate, far above Opus 4.7 (36%) and Gemini 3.1 Pro (50%). It's more likely to fabricate an answer than admit it doesn't know. Verify anything factual it gives you, especially in high-stakes domains.
Our Recommendation
Best for agent and shell work: GPT-5.5, because its 82.7% Terminal-Bench 2.0 score and "stays on task longer" behavior justify the premium for multi-step automation.
Best for reasoning and accuracy: Claude Opus 4.7, because it leads 6 of 10 shared benchmarks and hallucinates at 36% versus GPT-5.5's 86%.
Best value at frontier parity: Gemini 3.1 Pro, because it matches the index leaders for the lowest cost of the three at $2/$12.
Your next step: Run a 20-task pilot of GPT-5.5 on your actual agent workflow at "medium" effort. If your work is shell-driven and multi-step, the premium pays off. If it's reasoning or fact-heavy, route to Opus 4.7 instead and pocket the $5 output savings.
