Our production default caps the model's reasoning at 800 characters before it has to answer. Raising that cap to 1200 looked like a free win on paper. We ran the real test, with the bar for passing written down before we saw a single result, and it failed on latency, not on accuracy.
ModelBrain's optional reasoning pass asks the local model to think before it answers, inside a grammar that bounds how long that reasoning can run. The current default stops it at 800 characters. A model that's still mid-thought at 800 characters gets cut off, and a cut-off thought sometimes spills into the answer itself instead of a clean response. The obvious next question: does a longer leash just fix more of that, for free?
We tested it the way we test any change to the reasoning path: a fixed pass/fail bar, written down before the run, against the full 285-question evaluation set, on the real production prompt format, not a simplified one. No arm gets to pass by changing the definition of passing after we've seen its score.
We compared three versions side by side: the current 800-character default, the 1200-character candidate, and an unconstrained baseline with no reasoning cap at all. Each ran the complete 285-question set six separate times, with a different random seed each time, so a single lucky or unlucky run couldn't carry the result. That's 1,710 question-and-answer pairs per arm, 5,130 total.
Strict accuracy requires an exact match to the gold answer; lenient allows a correct answer embedded in extra text. Full methodology below.
The 1200-character version cut truncated reasoning by roughly two-thirds compared to the 800-character default, which was the headline number we were hoping to move. On one specific question we track because an earlier run truncated badly on it, accuracy did improve: 5 of 6 seeds landed on the right answer versus 3 of 6 at 800 characters.
But accuracy overall didn't move. Strict and lenient answer rates for the 1200-character arm came back statistically indistinguishable from the 800-character default, both inside the noise floor we'd measured from six baseline runs beforehand. More room to think stopped more reasoning from getting cut off mid-sentence, but it didn't translate into more right answers.
We set one hard bound going in, with no waiver written into the rule: response time at the 95th percentile had to stay within 20% of the baseline's. It's a latency promise to anyone running this on their own hardware, not a soft target to negotiate after the fact.
The 1200-character arm missed it in 4 of 6 runs, once by as much as 45%. The completion budget is identical between the 800 and 1200 character versions, and the longer arm wasn't retrying more often either, we checked. It simply used more of that shared budget on a given question more often, because there was more room to fill. More thinking time costs real wall-clock time even when the token budget doesn't change, and that's true whether or not it produces a better answer.
Our own rule, decided before we ran anything: a hard bound that fails stops the candidate regardless of what else looks good. So the 800-character default stays. The 1200-character result is recorded as the measured ceiling of this particular lever, not left as an open question to revisit without new evidence.
Before the run, we wrote down a prediction for one specific question in the set, a comparison between two board games where the gold answer competes against a cluster of plausible, similarly-dated distractors. If an 800-character cutoff was forcing a wrong guess, more room should fix it. If the question itself was the problem, more room would just let the model reason its way more confidently into the wrong answer.
It was the question. Zero of six 1200-character runs got it right, worse than either shorter arm. Extra reasoning room let the model enumerate further into the distractor cluster and land on a more specific, more wrong answer. That's now flagged as a question that needs rewriting, not a reasoning limit to engineer around.
Everything above came from one evaluation harness run against the real production code path, not a simplified stand-in:
Everything shipped from this round, the fix for the degenerate-answer bug mentioned above included, stayed inside that evaluation set. The candidate that moved a number without meeting the bar doesn't ship; the one that did, already has.