DictionaryEvaluation See it in actionBenchmarks explained
Benchmark Saturation
When an exam becomes so easy for top AIs that everyone scores near 100% and it stops ranking them.
Definition
The point where top models all score near a benchmark's ceiling, so it can no longer tell them apart. Humanity's Last Exam was built partly because LLMs reached over 90% accuracy on popular benchmarks such as MMLU. The usual fix is a newer, harder, or private test.
Example
If five models all score 97 to 99% on a 200-question test, the gaps are inside the noise; switch to a harder benchmark.