Every AI memory tool claims to be accurate. Almost none of them say against what. A single number with no stated test conditions, no breakdown by failure type, and no distinction between a strict match and a person's read of “basically right” is not a measurement. It is a marketing line dressed up as one.
Ask a vendor for their accuracy number and you'll usually get one figure. We publish two, because they measure different things and collapsing them hides the gap. Across 807 questions, each run three times and averaged, ModelBrain's recall test comes back at 70% against a strict exact match. Scored the way a person reading the answer would judge it, the same run comes back at 86%.
Both numbers are real, and neither is the whole story on its own. Strict scoring punishes a correct answer phrased differently from the gold text. Lenient scoring is more forgiving of an answer that is technically right but padded with more than it needed to say. Publish only the higher one and you aren't lying, you're choosing the number that flatters you. We publish both because a number without its scoring method isn't a number you can act on.
Broken down by what the test actually asks, the easy cases are close to perfect. Recalling a single fact and declining to answer a question with no answer in memory both score 100%, and so do the prompt-injection cases, meaning nothing planted in stored notes was treated as an instruction. Corrected facts, where an earlier answer has since been superseded, score 94%. Finding a fact buried in a long pile of notes and telling apart two lookalike entries both land at 90%.
The category that is actually hard is connecting two separate notes to answer one question. That scores 79%, the lowest of anything we test, and we publish it in the same table as the 100%s rather than in a footnote, because a results page that only shows you its best numbers isn't a results page. It's a highlight reel. One more thing sits deliberately outside the number: noticing on its own that two of your notes contradict each other. That isn't built yet, so it isn't counted.
We wrote the test as a spec before most of what it measures existed, so it couldn't get written around whatever the product already happened to do. That decision is written up here. The five questions behind it (what happens when you change your mind, can it tell two similar things apart, does deletion actually delete, what does it do with silence, where's the evidence) aren't ours alone to ask. Run them against anything you're evaluating, including us.
We wrote both the system being tested and the test that scores it, which is why the harness and the full methodology ship with v1: so you can run it yourself instead of taking our word. Until then, three things shape how to read these numbers. The gold answers are labeled by people who know what the system is supposed to do. The noise testing pads real notes with synthetic distractor content, and synthetic noise behaves differently from the messy, half-updated notes a real vault accumulates over a year. And citation fidelity, whether the system's claimed source actually supports what it says, depends partly on which assistant model is reading the citation back to you, not on ModelBrain alone.
None of that changes the 86% or the 70%. It is why the five questions exist: run them against your own data, not just ours.
The full category breakdown is on the fact sheet today. The harness and complete methodology ship with v1. If you find a way the test flatters us that we haven't owned up to here, the fastest way to tell us is to run your own version and show us where it breaks: hello@modelbrain.net.