DictionaryEvaluation See it in actionBenchmarks explained
Benchmark
A standard exam for AI models: the same questions for every model, graded automatically, so you can compare scores fairly.
Definition
A fixed set of tasks with an automatic grader, run the same way on every model so results can be compared. The score is the share of tasks passed. Scores are only comparable within the same variant and setup (attempts, tools, effort, harness), and independent evaluators such as Epoch AI track results across models over time.
Example
Model A solves 410 of 500 SWE-bench Verified tasks, scoring 82%. Comparing that with another lab's pass@5 score would be misleading.
Appears in
- LongGenBench: Long-context Generation BenchmarkarXiv
- TRACE: TRajectory Attribution for Automated Context EngineeringarXiv
- Self-Evolving Coding AgentsarXiv
- SkillReducer: Optimizing LLM Agent Skills for Token EfficiencyarXiv
- Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic TasksarXiv
- DeepCode: Open Agentic CodingarXiv