GPT-5.6 Sol ARC-AGI-3 score triples via API settings
TECH

GPT-5.6 Sol ARC-AGI-3 score triples via API settings

23+
Signals

Strategic Overview

  • 01.
    OpenAI reports that enabling two settings in its Responses API - Retained Reasoning and Compaction - tripled GPT-5.6 Sol's score on the ARC-AGI-3 public task set, while using roughly six times fewer output tokens per game.
  • 02.
    With the official ARC-AGI-3 harness, GPT-5.6 Sol scored 13.3% on the public task set; using OpenAI's Responses API with retained reasoning and compaction turned on, the same model scored 38.3%.
  • 03.
    GPT-5.6 Sol's officially verified score under the standardized ARC Prize harness is 7.78% (Max reasoning effort) on the public set; it was the first model to win an ARC-AGI-3 public game.
  • 04.
    OpenAI claims that with retained reasoning and compaction enabled, GPT-5.6 Sol's 38.3% score surpasses Claude Opus 5's official ARC-AGI-3 score of 30.2%.
  • 05.
    Claude Opus 5's officially verified ARC-AGI-3 score is 30.16% (High reasoning effort), the highest official score recorded on the benchmark to date, and it completed five environments no prior model had beaten.
  • 06.
    In the official, non-adjusted test harness, GPT-5.6 Sol scored 7.8% while the prior-generation GPT-5.5 scored just 0.4%.

What Retained Reasoning and Compaction Actually Do

The official ARC-AGI-3 evaluation wipes a model's working memory after every single move: it discards the model's private reasoning and, once the running history exceeds the context window, drops earlier actions too, forcing the model to reorient in the game environment from scratch on each turn[2]. OpenAI's fix was two Responses API settings: Retained Reasoning, which keeps chain-of-thought persisting across turns when a developer passes along the previous response ID, and Compaction, which summarizes older context instead of truncating or discarding it[1]. With both switched on, GPT-5.6 Sol's ARC-AGI-3 public-set score rose from an official baseline in the 7.8-13.3% range to 38.3%, while output tokens per game fell roughly sixfold[1][2].

A Benchmark Built to Prevent This Exact Kind of Tweak

ARC-AGI-3 discards reasoning between moves specifically so a score reflects a model's own problem-solving, not accumulated context or harness engineering[2]. OpenAI's settings reintroduce exactly the continuity the standardized evaluation was built to strip out, then use the resulting score to argue GPT-5.6 Sol edges past Claude Opus 5. ARC Prize co-founder Francois Chollet's response drew a distinction between an unfair 'custom-made' harness and a legitimate 'general-purpose API setting...available to all API users' - but he conditioned that acceptance on the settings and cost being clearly disclosed[3]. That the acceptance needed a disclosure caveat at all is itself a sign the two evaluations are not measuring the same thing.

The Leapfrogging That Already Happened

Timing matters here. GPT-5.6 Sol's officially verified ARC-AGI-3 score was 7.78% under Max reasoning effort[4]. Only days before OpenAI's settings post, Claude Opus 5 set a new official record of 30.16% (30.2%) under High effort, verified by ARC Prize as the highest score on the benchmark to date[5]. OpenAI's 38.3% figure, published after that record, is framed as beating Opus 5 - but Opus 5 was never run with the same retained-reasoning-plus-compaction configuration, so the comparison sets one model's custom-tuned harness against another model's standardized one[3].

What This Says About Reading AI Benchmarks Going Forward

The same model weights produced scores nearly five times apart - 7.8% to 13.3% under the standard harness versus 38.3% under OpenAI's own settings - which means harness and API configuration can matter as much as which model is being tested[1][6]. As one outlet summarized the underlying dynamic: 'Evaluations rarely test models in isolation. They are also influenced by API settings and harness design.'[6]Any future lab claim of 'model X beats model Y' on a benchmark is worth reading as a claim about model plus harness plus settings, not the model alone.

Historical Context

2026-07
GPT-5.6 series (Sol, Terra, Luna) launched, with Sol positioned as OpenAI's flagship reasoning and agentic model.
2026-07-27
Claude Opus 5 set a new official ARC-AGI-3 record at 30.2%, nearly four times the previous official record of 7.8% set by GPT-5.6 Sol under standardized evaluation; the ARC Prize verified the result.
N/A
ARC-AGI-3 is the third-generation Abstraction and Reasoning Corpus benchmark, an interactive-reasoning evaluation designed to measure fluid, novel problem-solving ability that remains far from saturated for frontier models.

Power Map

Key Players
Subject

GPT-5.6 Sol ARC-AGI-3 score triples via API settings

OP

OpenAI

Model developer; published the blog post claiming the tripled score and the comparison to Claude Opus 5 using its own Responses API settings

AN

Anthropic

Developer of Claude Opus 5, which holds the official ARC-AGI-3 record (30.2%) that OpenAI's non-standard result claims to beat

AR

ARC Prize Foundation / Francois Chollet

Runs the official, standardized ARC-AGI-3 benchmark and evaluation harness; ARC Prize co-founder Chollet publicly commented on the distinction between acceptable general-purpose API settings and unfair custom harnesses

Fact Check

6 cited
  1. [1] OpenAI Says GPT 5.6's Score On ARC-AGI 3 Tripled After Turning On Two API Settings
  2. [2] OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
  3. [3] OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness
  4. [4] GPT-5.6 Sol Results
  5. [5] Claude Opus 5 Results
  6. [6] OpenAI Triples Benchmark Scores With Simple Settings

Source Articles

Top 4

THE SIGNAL.

Analysts

Chollet distinguished between unfair 'custom-made' harnesses and acceptable 'general-purpose API settings...available to all API users,' while conditioning that acceptance on the settings and cost being clearly disclosed.

Francois Chollet (ARC Prize co-founder)
Distinguishes between unfair custom-made benchmark harnesses and legitimate general-purpose API settings
The Crowd

OpenAI says two API settings raised GPT-5.6 Sol’s ARC-AGI-3 public-set score from 13.3% to 38.3%. And this happened because, the ARC-AGI official evaluation kept wiping 5.6 Sol's working memory. i.e. it discarded private reasoning after every move and later removed older actions once the running history exceeded its window.

@@rohanpaul_ai44

Is GPT‑5.6 Sol now better than Opus 5 on ARC‑AGI‑3? Short answer: not on the official leaderboard, but the comparison is more complicated than it first appeared. let me explain, because there is a bit of confusion. ARC‑AGI‑3 tests whether models can learn unfamiliar 2D games through trial and error...

@@kimmonismus440

Opus 5 just jumped the ARC-AGI-3 score from 7.8% to 30.2% The previous score was GPT-5.6 Sol, which came out..... two weeks ago.

@@BenjaminDEKR35

Tibo addresses 5.6 Sol Score on ARC-AGI-3

@u/Apple_macOS272
Broadcast
Ep 126: OpenAI's GPT-5.6 Sol just tripled its ARC-AGI-3 score using two API settings that also cut output tokens by six times

Ep 126: OpenAI's GPT-5.6 Sol just tripled its ARC-AGI-3 score using two API settings that also cut output tokens by six times

OpenAI is so back... GPT 5.6 Sol first look

OpenAI is so back... GPT 5.6 Sol first look

GPT 5.6 SOL y NUEVO KIMI 3 ¡La IA AGÉNTICA ya es REAL!

GPT 5.6 SOL y NUEVO KIMI 3 ¡La IA AGÉNTICA ya es REAL!