@tokalator4.2.0…installs…downloads
All benchmarks
Expert knowledge

GPQA Diamond

PhD-level science questions that can't be simply searched online.

Tasks
198
Released
Nov 2023
Made by
New York University, Cohere and Anthropic
Score
Percent answered correctly
In plain words

Graduate-school biology, physics and chemistry questions so tricky that smart non-experts with Google still mostly get them wrong.

How it is marked

The chosen option is compared with the correct answer.

Humans: PhD-level experts score 65% on the wider 448-question set; skilled non-experts with web access score 34% source

Scores we track

Each number is the provider's own published result. Tap one for its source.

  1. 1GPT-6 Astra96%
  2. 2GPT-5.6 Sol94.6%
  3. 3Gemini 3.1 Pro94.3%
  4. 4Claude Opus 4.794.2%
  5. 5Claude Mythos 594.1%
  6. 6Claude Opus 4.893.6%
  7. 7GPT-5.593.6%
  8. 8GPT-5.492.8%
  9. 9GPT-5.3-Codex92.6%
  10. 10Claude Opus 4.691.3%
  11. 11Gemini 3 Flash90.4%
  12. 12Claude Sonnet 4.689.9%
  13. 13GPT-5.4 mini88%
  14. 14Claude Opus 4.587%
  15. 15Gemini 2.5 Pro86.4%
  16. 16GPT-5.4 nano82.8%
  17. 17Claude Haiku 4.573%

Why it matters

It checks whether AI can reason like a scientist, not just recall facts it could look up.

Good to know

  • The Diamond set keeps only questions where both expert checkers agreed and most non-experts failed.
  • With only 198 questions, a few answers can move the score noticeably.
  • The authors ask people not to post the questions online, to keep them out of AI training data.