@tokalator4.2.0…installs…downloads
All benchmarks
Long memory

OpenAI MRCR

Pick the right repeated item from a very long chat.

Tasks
2,400
Released
Apr 2025
Made by
OpenAI, building on a Google DeepMind test from the Gemini team
Score
Average text-match score
In plain words

Like scrolling back through a months-long chat to find the second of several near-identical poems you once asked for.

Example taskSource

In a long chat with several poems about tapirs, return the second tapir poem exactly.

A long made-up conversation where similar requests repeat 2, 4 or 8 times. The model must return one specific earlier reply.

How it is marked

The answer must start with a given code, then it is scored by how closely its text matches the correct reply.

Why it matters

Huge context windows only help if the model can find the right detail inside them.

Good to know

  • Conversations range from about 4,000 to over 1 million tokens; scores depend heavily on length.
  • Results are often reported for one setting, such as 8 repeats, so check which one.