All benchmarksExpert knowledge In plain words Example taskSource
Humanity's Last Exam
Very hard expert questions across over a hundred subjects.
- Tasks
- 2,500
- Released
- Jan 2025
- Made by
- Center for AI Safety and Scale AI
- Score
- Percent answered correctly; calibration error is also reported
A final exam written by about a thousand professors and researchers, using only questions the best AI models got wrong.
Translate a Roman-era tombstone inscription written in Palmyrene script, given a letter-by-letter transcription.
One question with a single clear answer. About 24% are multiple choice; about 14% include an image.How it is marked
An AI judge compares the final answer with the known correct answer, allowing equivalent forms like fractions versus decimals.
Scores we track
Each number is the provider's own published result. Tap one for its source.
- 1Claude Opus 5.564.4%
- 2Claude Fable 5.160.9%
- 3Claude Mythos 5.160.9%
- 4Claude Mythos 559%
- 5Claude Sonnet 5.556.9%
- 6Claude Opus 556.3%
- 7Claude Opus 4.849.8%
- 8Claude Opus 4.746.9%
- 9Claude Haiku 5.545.9%
- 10Gemini 3.1 Pro44.4%
- 11Claude Sonnet 543.2%
- 12GPT-5.541.4%
- 13Claude Opus 4.640%
- 14GPT-5.439.8%
- 15Gemini 3 Flash33.7%
- 16Claude Sonnet 4.633.2%
- 17Claude Opus 4.530.8%
- 18GPT-5.4 mini28.2%
- 19GPT-5.4 nano24.3%
- 20Gemini 2.5 Pro21.6%
Why it matters
Older knowledge tests are nearly maxed out. This one still shows how far AI is from top human experts.
Good to know
- Questions were kept only if leading AI models failed them, so early low scores are partly by design.
- A private set of questions is held back to catch models trained on the public ones.
- The paper's abstract says 'dozens of subjects' while its main text and site say over a hundred.