DictionaryEvaluation See it in actionBenchmarks explained
Benchmark Contamination
When an AI has already seen the test answers during training, so its high score says less about real ability.
Definition
When benchmark questions or answers leak into a model's training data, so it can score well by recall rather than skill. A 2024 survey defines benchmark data contamination as models inadvertently incorporating evaluation benchmark information from their training data, leading to inaccurate or unreliable evaluation. Common responses are private or held-out test sets, freshly written tasks, and tasks dated after a model's training cutoff.
Example
A model aces GitHub issues filed before its training cutoff but drops sharply on issues filed after it, a classic contamination signal.