All benchmarksLong memory In plain words Example taskSource
OpenAI MRCR
Pick the right repeated item from a very long chat.
- Tasks
- 2,400
- Released
- Apr 2025
- Made by
- OpenAI, building on a Google DeepMind test from the Gemini team
- Score
- Average text-match score
Like scrolling back through a months-long chat to find the second of several near-identical poems you once asked for.
In a long chat with several poems about tapirs, return the second tapir poem exactly.
A long made-up conversation where similar requests repeat 2, 4 or 8 times. The model must return one specific earlier reply.How it is marked
The answer must start with a given code, then it is scored by how closely its text matches the correct reply.
Why it matters
Huge context windows only help if the model can find the right detail inside them.
Good to know
- Conversations range from about 4,000 to over 1 million tokens; scores depend heavily on length.
- Results are often reported for one setting, such as 8 repeats, so check which one.