Docs / Reasoning mode

Reasoning mode, explained

Plain recall ranks your notes and hands back the closest matches; that's the default, and it's what most questions need. Reasoning mode is an optional extra pass on the same call that reads those notes and synthesizes one checked answer, instead of leaving you to read the matches yourself.

Plain search vs. reasoning

Every recall call does the ordinary ranked search first, and always returns those hits (results) — that part never changes. Add reasoning: true and, after that same search, ModelBrain also sends the matching notes through a local reasoning model to work out one answer across them, verifies that answer actually follows from the notes it cites, and returns it as answer only once that check passes.

"What's the launch date for the project Priya leads?" needs reasoning, since the answer isn't in any single note by itself. "Find my note on Falcon" doesn't — plain search already returns it, faster.

What it costs you

Reasoning runs a real model on your own machine, not a network call, so it's meaningfully slower than plain search: a multi-second to tens-of-seconds local model call, more on CPU-only hardware. It also fails closed rather than ever guessing — if it can't produce an answer it can verify, recall returns reasoning_unavailable with a specific reason instead, and your plain search results are still there either way.

Two optional add-ons
widen

Before reasoning runs, generates a few alternative phrasings of your question, pulls in each result's directly linked memories one hop out, and reasons over that wider pool instead of just the first search's matches. Useful when the answer is scattered across notes that don't share your exact wording. Does nothing unless reasoning is also on.

verify

Runs the whole reasoning pass twice, independently, and only returns an answer if both runs reach the same conclusion — disagreement is treated as a failed check, not a coin flip. A real latency cost (up to 14 model calls in the worst case, against 4 for a single pass) for a second, independent check. Does nothing unless reasoning is also on.

Checking before you pay the latency

Every recall response, whether or not you asked for reasoning, includes a status: whether a reasoning pass would actually run right now, which tier is selected, and whether that tier's model is ready, still downloading, or failed to load. Check that first if you want to know whether asking for reasoning: true on your next call is worth the wait.

When reasoning can't answer

A reasoning call either returns a verified answer, or reasoning_unavailable naming exactly why — never a guess it couldn't check.

reason
Meaning
model_provisioning
The reasoning model is still downloading (one-time).
hardware_too_slow
This machine is below even the lightest tier's requirements.
model_load_failed
The model file failed to load.
timeout
No verified answer inside the time budget — the model was still working, it just ran out of wall-clock time.
call_budget_exceeded
The model ran out of reasoning steps before reaching a verified answer — a different limit than timeout, counted in steps rather than time.
verification_failed
An answer was generated but failed its own citation check (or, with verify on, the two independent passes disagreed), so it was withheld.
busy
Another reasoning call is already in flight on this connection; only one runs at a time.
Which model runs, and choosing it yourself

ModelBrain measures your machine's free RAM and CPU cores and picks a model tier for you automatically; see hardware requirements & model tiers for exactly what each tier needs and downloads. To pin a tier yourself instead of auto-selecting:

modelbrain_admin reasoning set-tier <auto|minimal|default|high|off>

A running daemon picks up the change within about 30 seconds, no restart needed. off disables reasoning entirely on that machine; plain search, remembering and recall keep working regardless of tier.

Questions

What is reasoning mode?

An optional pass on recall (reasoning: true) that runs after the ordinary ranked search. It sends the matching notes through a local reasoning model to synthesize one answer, checks that answer against the notes it cites, and returns it only if the check passes. The plain ranked hits are always returned too, whether or not reasoning is requested.

Why is reasoning slower than plain recall?

It runs a real local model instead of just ranking existing embeddings, so it takes a multi-second to tens-of-seconds local model call rather than a near-instant vector search, more on CPU-only hardware.

Does reasoning mode ever make something up?

It's built to fail closed instead. Every reasoning answer is checked against the notes it cites before it's returned; if that check fails, or the pass can't finish for another reason, recall returns reasoning_unavailable with a specific reason instead of a guess.

How do I pick which tier reasoning uses?

It auto-selects based on your machine's RAM and CPU cores. Override it with modelbrain_admin reasoning set-tier <auto|minimal|default|high|off>; a running daemon picks it up within about 30 seconds, no restart.

Related

What each tier needs, and what it downloads.

Hardware requirements & model tiers →
Full parameter reference

Every recall field, plus the complete tool reference.

MCP tools reference →