@tokalator4.2.0…installs…downloads
All benchmarks
Puzzle reasoning

ARC-AGI-2

Colored-grid puzzles that test learning a new rule fast.

Tasks
120
Released
Mar 2025
Made by
ARC Prize Foundation
Score
Percent of tasks solved within two tries, reported alongside cost per task
In plain words

Like an IQ-test puzzle: see a few before-and-after picture pairs, work out the hidden rule, then apply it to a new picture.

Example taskSource

Colored rectangles with holes act as a key: each shape gets the color whose key has the same number of holes.

A few example pairs of colored grids, then a new grid. The solver must produce the correct output grid.

How it is marked

The output grid must match the answer exactly. Two attempts are allowed per task.

Humans: Human panel average 60%; every task was solved by at least two people within two tries source

Scores we track

Each number is the provider's own published result. Tap one for its source.

  1. 1GPT-6 Astra95%
  2. 2GPT-5.6 Sol92.5%
  3. 3Claude Opus 590.42%
  4. 4Claude Fable 5.190%
  5. 5Claude Mythos 5.190%
  6. 6Claude Fable 589.2%
  7. 7Claude Mythos 589.2%
  8. 8GPT-5.585%
  9. 9Gemini 3.1 Pro77.1%
  10. 10GPT-5.473.3%
  11. 11Gemini 3.5 Flash72.1%
  12. 12Claude Opus 4.668.8%
  13. 13Claude Sonnet 4.658.3%
  14. 14Claude Opus 4.537.6%
  15. 15Gemini 3 Flash33.6%
  16. 16Gemini 2.5 Pro4.9%

Why it matters

It targets adapting to something truly new, which ordinary people do easily but AI still finds hard.

Good to know

  • The 120 count is per evaluation set; there are public, semi-private and private sets of 120 each.
  • Scores should be read with cost, since spending more computing power can raise results.
  • The paper reports 66% of human attempts succeeded, versus 60% on the announcement page.