@tokalator4.2.0…installs…downloads
All benchmarks
Expert knowledge

Humanity's Last Exam

Very hard expert questions across over a hundred subjects.

Tasks
2,500
Released
Jan 2025
Made by
Center for AI Safety and Scale AI
Score
Percent answered correctly; calibration error is also reported
In plain words

A final exam written by about a thousand professors and researchers, using only questions the best AI models got wrong.

Example taskSource

Translate a Roman-era tombstone inscription written in Palmyrene script, given a letter-by-letter transcription.

One question with a single clear answer. About 24% are multiple choice; about 14% include an image.

How it is marked

An AI judge compares the final answer with the known correct answer, allowing equivalent forms like fractions versus decimals.

Scores we track

Each number is the provider's own published result. Tap one for its source.

  1. 1Claude Opus 5.564.4%
  2. 2Claude Fable 5.160.9%
  3. 3Claude Mythos 5.160.9%
  4. 4Claude Mythos 559%
  5. 5Claude Sonnet 5.556.9%
  6. 6Claude Opus 556.3%
  7. 7Claude Opus 4.849.8%
  8. 8Claude Opus 4.746.9%
  9. 9Claude Haiku 5.545.9%
  10. 10Gemini 3.1 Pro44.4%
  11. 11Claude Sonnet 543.2%
  12. 12GPT-5.541.4%
  13. 13Claude Opus 4.640%
  14. 14GPT-5.439.8%
  15. 15Gemini 3 Flash33.7%
  16. 16Claude Sonnet 4.633.2%
  17. 17Claude Opus 4.530.8%
  18. 18GPT-5.4 mini28.2%
  19. 19GPT-5.4 nano24.3%
  20. 20Gemini 2.5 Pro21.6%

Why it matters

Older knowledge tests are nearly maxed out. This one still shows how far AI is from top human experts.

Good to know

  • Questions were kept only if leading AI models failed them, so early low scores are partly by design.
  • A private set of questions is held back to catch models trained on the public ones.
  • The paper's abstract says 'dozens of subjects' while its main text and site say over a hundred.