All benchmarksPuzzle reasoning In plain words Example taskSource
ARC-AGI-2
Colored-grid puzzles that test learning a new rule fast.
- Tasks
- 120
- Released
- Mar 2025
- Made by
- ARC Prize Foundation
- Score
- Percent of tasks solved within two tries, reported alongside cost per task
Like an IQ-test puzzle: see a few before-and-after picture pairs, work out the hidden rule, then apply it to a new picture.
Colored rectangles with holes act as a key: each shape gets the color whose key has the same number of holes.
A few example pairs of colored grids, then a new grid. The solver must produce the correct output grid.How it is marked
The output grid must match the answer exactly. Two attempts are allowed per task.
Humans: Human panel average 60%; every task was solved by at least two people within two tries source
Scores we track
Each number is the provider's own published result. Tap one for its source.
- 1GPT-6 Astra95%
- 2GPT-5.6 Sol92.5%
- 3Claude Opus 590.42%
- 4Claude Fable 5.190%
- 5Claude Mythos 5.190%
- 6Claude Fable 589.2%
- 7Claude Mythos 589.2%
- 8GPT-5.585%
- 9Gemini 3.1 Pro77.1%
- 10GPT-5.473.3%
- 11Gemini 3.5 Flash72.1%
- 12Claude Opus 4.668.8%
- 13Claude Sonnet 4.658.3%
- 14Claude Opus 4.537.6%
- 15Gemini 3 Flash33.6%
- 16Gemini 2.5 Pro4.9%
Why it matters
It targets adapting to something truly new, which ordinary people do easily but AI still finds hard.
Good to know
- The 120 count is per evaluation set; there are public, semi-private and private sets of 120 each.
- Scores should be read with cost, since spending more computing power can raise results.
- The paper reports 66% of human attempts succeeded, versus 60% on the announcement page.