Home / Recall test results
ModelBrain recall test results
Across 807 questions, run three times and averaged, ModelBrain gave the right answer 86% of the time as a person reading it would judge, and 70% under a strict exact match. Every case is below, including the weakest one.
The two headline numbers
Scoring
Result, 807 questions, 3 runs averaged
Lenient: a person would call it right
86%
Strict: exact match to the gold answer
70%
We publish both because they measure different things. Strict scoring punishes a correct answer phrased differently from the gold text. Lenient scoring forgives a correct answer padded with extra words. Why we show both: what “accurate” should mean for an AI's memory.
Right-answer rate by test case
Test case
Right answer rate
A single fact you told it once
100%
A prompt-injection attempt in your notes
100%
A question with no answer in memory (correctly declines)
100%
A fact you've since corrected
94%
A long pile of notes, right one buried
90%
Lookalike facts designed to confuse it
90%
Connecting two separate notes to answer one question
79%
The easy cases are close to perfect. The hard one is connecting two separate notes to answer one question, at 79%, and we put it in the same table as the 100% results.
What each case checks
- A single fact you told it once. Save one fact, ask for it later in a new conversation.
- A prompt-injection attempt. A note contains an instruction aimed at the assistant. Passing means nothing planted in stored notes was treated as an instruction.
- No answer in memory. Passing means it says so instead of guessing.
- A corrected fact. An earlier answer has since been superseded; it should give the new one.
- A long pile of notes. The right note is buried among many others.
- Lookalike facts. Two similar projects or entries that are easy to confuse.
- Two notes, one question. Neither note answers it alone.
A second, larger look: six seeds on the 285-question set
When we tested whether a longer reasoning ceiling was worth shipping, we ran the production configuration on a fixed 285-question evaluation set six times with different random seeds. This is a different set from the 807-question run above, so the numbers differ a little.
Production default, mean of 6 seeds
Result
Strict answer accuracy
68.4%
Lenient answer accuracy
85.1%
The full comparison, including the candidate we rejected on latency, is in why we didn't ship a 1200-character reasoning ceiling.
What is not counted, and what this does not prove
- Contradiction detection is not counted. Noticing on its own that two of your notes disagree is not built yet, so it is outside the number.
- We wrote both the system and the test. The harness and full methodology ship with v1 so you can run it yourself.
- Gold answers are labeled by people who know what the system should do.
- The noise tests use synthetic distractors, which behave differently from a real vault's half-updated notes after a year.
- Citation fidelity depends partly on which assistant model reads the citation back to you, not on ModelBrain alone.
The test was written before most of what it measures existed. How and why, and five questions to ask any AI memory tool, including this one.