GPT-5.6 Explained: Sol, Terra, and Luna — The Model Family That Cleared a White House Review

OpenAI's new model family, GPT-5.6, became generally available across ChatGPT, the API, Codex, and GitHub Copilot on July 9, 2026. But this launch's backstory is far more interesting than a typical release: it became the first AI model required to clear a White House review before it could ship at all.

OpenAI


Government First, Public Second

GPT-5.6's story starts on June 26, 2026. That day, OpenAI launched the model as a limited preview restricted to government-vetted partners only. The request came from the White House's Office of the National Cyber Director and Office of Science and Technology Policy, citing the model's advanced cybersecurity capabilities as the reason.

It marked the first time the US government had asked an American AI company — on a nominally voluntary basis — to restrict a model's launch before it shipped publicly. After that 12-day closed-door testing period, the model went generally available on July 9.

AI-Powered Mobile App Development
AI-Powered Mobile App Development
Learn AI-powered mobile application development techniques.
Go to course →

Three Models, Three Different Needs

GPT-5.6 didn't ship as a single model — it arrived as a family with three distinct tiers:

  • Sol — The flagship. It delivers the highest performance and can hit 750 tokens/second on Cerebras hardware. Built for the heaviest reasoning and coding workloads.
  • Terra — A balanced mid-tier model for everyday use, delivering solid performance for most workflows at a reasonable cost.
  • Luna — The cost-efficient, lightweight option, optimized for high-volume, simple tasks.

This three-tier structure is a clear signal that OpenAI has moved away from a single "best model" strategy and toward a portfolio approach that lets users choose price and performance based on their actual use case.


The Benchmarks: How Good Is Sol, Really?

OpenAI's claims are strong, but the benchmark picture is a bit more nuanced than the headline numbers suggest:

  • On the Artificial Analysis Coding Agent Index, Sol set a new state-of-the-art score of 80 at max reasoning effort — 2.8 points above Claude Fable 5.
  • But on SWE-bench Pro, Sol scored only 64.6%, trailing other frontier models released around the same time. Sol isn't the leader on every benchmark — there's a real gap between where it excels and where it lags.
  • ARC-AGI testing produced an interesting story of its own: right after launch, Sol underperformed expectations on this visual-reasoning benchmark. OpenAI's research team dug in and found the harness it ran on wasn't letting the model retain what it had already learned mid-task. By enabling two API settings — retained reasoning and context compaction — they tripled the scores while using 6x fewer output tokens. After that fix, Sol reached 29.3% on ARC-AGI-3 and 92.5% on ARC-AGI-2, and became the first model to win a public ARC-AGI-3 game outright.

That detail carries a broader lesson: for today's frontier models, the harness and API configuration running underneath can matter almost as much as the raw model itself in determining final performance.


Why This Matters

GPT-5.6's government review process was one of the early links in a pattern that repeated throughout the summer. Just a few weeks earlier, Anthropic's Claude Fable 5 had gone offline for 18 days over export-control rules; now OpenAI was going through a comparable review before it could even launch its model. Together, these two episodes point to a new industry norm: once a model crosses a certain capability threshold, government review is now routine.

With deep integration into ChatGPT, the API, Codex, and GitHub Copilot, GPT-5.6 reached a massive user base immediately — and, launching the same week as Meta's Muse Spark 1.1, it helped make this one of the busiest model-release weeks of summer 2026.

GPT-5.6OpenAIChatGPTCodexGitHub CopilotAI NewsLLM
Tuncer Bağçabaşı
Tuncer Bağçabaşı
Software Engineer & AI Researcher
← All posts