All benchmarksExpert knowledge In plain words
GPQA Diamond
PhD-level science questions that can't be simply searched online.
- Tasks
- 198
- Released
- Nov 2023
- Made by
- New York University, Cohere and Anthropic
- Score
- Percent answered correctly
Graduate-school biology, physics and chemistry questions so tricky that smart non-experts with Google still mostly get them wrong.
How it is marked
The chosen option is compared with the correct answer.
Humans: PhD-level experts score 65% on the wider 448-question set; skilled non-experts with web access score 34% source
Scores we track
Each number is the provider's own published result. Tap one for its source.
- 1GPT-6 Astra96%
- 2GPT-5.6 Sol94.6%
- 3Gemini 3.1 Pro94.3%
- 4Claude Opus 4.794.2%
- 5Claude Mythos 594.1%
- 6Claude Opus 4.893.6%
- 7GPT-5.593.6%
- 8GPT-5.492.8%
- 9GPT-5.3-Codex92.6%
- 10Claude Opus 4.691.3%
- 11Gemini 3 Flash90.4%
- 12Claude Sonnet 4.689.9%
- 13GPT-5.4 mini88%
- 14Claude Opus 4.587%
- 15Gemini 2.5 Pro86.4%
- 16GPT-5.4 nano82.8%
- 17Claude Haiku 4.573%
Why it matters
It checks whether AI can reason like a scientist, not just recall facts it could look up.
Good to know
- The Diamond set keeps only questions where both expert checkers agreed and most non-experts failed.
- With only 198 questions, a few answers can move the score noticeably.
- The authors ask people not to post the questions online, to keep them out of AI training data.