Paying your AI to think longer doesn't fix its mistakes. Jev still catches more of them.
Jamie Watters
Operational resilience and AI delivery practitioner. Technology since 1985.

Who this is for. Anyone deciding whether to give a model room to reason before it hands back a structured verdict, because that room costs real time and money and the obvious hope is that it closes an accuracy gap. By the end you'll know whether it did, on the one hard task I can measure it on, and how to tell a real shift from two claims of luck.
Skip it if you want the confidence-threshold work: that's the piece before this one, and this one doesn't repeat it. This is also not a Jev tutorial.
On the length. About twelve minutes. The result itself is short. Most of the length is showing my working on the parts of it that don't fully support the headline, because a benchmark you can't argue with is a benchmark you can't check.
Two pieces ago I said the biggest open question about Jev's lead on the hard 18 was whether the frontier models had been handicapped: forced straight into a JSON verdict, no room to work through a claim before answering it. Every reviewer raised it independently. I said the run was next and I'd publish it either way.
It ran on 21 September. GPT-5.4, Claude Sonnet 5 and Gemini 3.1 Pro all got a scratchpad this time: think in plain prose first, then answer into the same structured verdict as before. None of the three improved on the hard 18. Jev's lead over the best of them widened, from 11.1 points to 16.7. It cost 2.4 to 4.2 times more per claim to find that out.
That's the headline, and it isn't the whole finding. The drops aren't statistically significant on eighteen items, so the honest sentence is that reasoning did not help, not that it made the models worse. And two things inside this run cut the other way. On two of the three models, the verdict call barely does anything once the scratchpad exists. And on the claims where reasoning did change an answer, the model's objection was often right about the passage. That's a finding against how I built the benchmark, not against the models, and it turned out to be the most useful thing this run produced.
What the reasoning arm actually did
Two calls per claim instead of one, added to JevBench as a new arm. The first gives the model the claim, the passage, and the same three label definitions the schema-only arm used, then asks it to think in plain prose, under 200 words, and not to answer yet. The second asks for the verdict in the same JSON schema as before, with that scratchpad already sitting in the conversation.
Two calls rather than one relaxed call, so the reasoning step's cost and time are measurable on their own instead of buried inside a single number. The scratchpad prompt says nothing the schema-only prompt didn't already say: no label flagged as likely, nothing about how the labels are distributed, nothing about what kind of errors this set is built from.
Same 42 claims as the first two pieces. Jev was not re-run: its row below is the 20 September run, unchanged.
The numbers
| all 42 | the easy 24 | the hard 18 | cost per claim | seconds | |
|---|---|---|---|---|---|
| Jev, not re-run | 85.7% | 91.7% | 77.8% | $0.000025 | 0.23 |
| GPT-5.4, schema only | 85.7% | 100% | 66.7% | $0.001289 | 1.01 |
| GPT-5.4, reasoning allowed | 83.3% | 100% | 61.1% | $0.005450 | 3.77 |
| Sonnet 5, schema only | 85.7% | 100% | 66.7% | $0.002228 | 2.76 |
| Sonnet 5, reasoning allowed | 76.2% | 100% | 44.4% | $0.006439 | 6.23 |
| Gemini 3.1 Pro, schema only | 81.0% | 100% | 55.6% | $0.004824 | 3.78 |
| Gemini 3.1 Pro, reasoning allowed | 78.6% | 95.8% | 55.6% | $0.011545 | 8.71 |
Read the hard column on JevBench's own numbers and none of the three frontier models improved. GPT-5.4 goes from 12 of 18 to 11. Sonnet 5 goes from 12 of 18 to 8. Gemini 3.1 Pro stays at 10 of 18, the same ten claims right both times. Jev, not re-run, still gets 14 of 18.
Against GPT-5.4's reasoning-allowed 11 of 18, that's a 16.7-point JevBench lead, three claims. Against the best either arm managed, the 12 of 18 that GPT-5.4 and Sonnet 5 both reached schema-only, it's 11.1 points, two claims. The lead didn't shrink. It grew, because the frontier models had more room to move than Jev did, and where they moved, more of them moved down.
What the reasoning arm cost to find that out: 4.23 times the price and 3.74 times the latency for GPT-5.4, 2.89 times and 2.26 times for Sonnet 5, 2.39 times and 2.31 times for Gemini. Against Jev, the cheapest reasoning-allowed arm still runs at roughly 220 times the cost per claim.

The part I can't claim
Eighteen items is not enough to say any of that with confidence, and I said so in the piece before this one. It applies harder here. Exact McNemar on the discordant pairs: GPT-5.4, 2 rows right only schema-only, 1 right only with reasoning, p = 1.000. Sonnet 5, 5 rows right only schema-only, 1 right only with reasoning, p = 0.219. Gemini 3.1 Pro, 0 and 0, identical answers on all 18, p = 1.000.
None of that clears even a loose bar for significance. So "reasoning made the frontier models worse" is not a sentence this run supports. What it supports is narrower, and it still answers the question that was asked: reasoning did not close the gap, for any of the three, and it did not produce a single net gain on the hard set anywhere.
Run the same test the other way, Jev against each reasoning-allowed model, and it favours Jev every time it's asked. But only the Sonnet 5 comparison clears the usual 0.05 bar, at p = 0.031, and it wouldn't survive being checked three times over. GPT-5.4 is p = 0.250, Gemini p = 0.125.
Where the answers moved
Nine verdicts changed on the hard set, seven of them from right to wrong, two from wrong to right.
Five of the nine moved towards the harsher reading, calling a claim unsupported when the arm
before it hadn't, and all five of those were wrong to move there.
Look at what happened in the model's own words rather than just the label, and something more
interesting than "the models got confused" shows up. Take rz-05, a claim built from a paper's
author affiliations: "Tencent AI Seattle Lab with Washington University, Maryland and UT Dallas."
Sonnet 5, straight into the schema, called that supported and was right to, at 0.92 confidence.
Given a scratchpad first, it read the actual affiliation list. Four separate institutions, it found,
including Washington University in St Louis and University of Maryland as two different places. So
it called the sentence unsupported: "the claim's wording incorrectly implies Washington University
is located in or associated with Maryland, when in fact Washington University is in St. Louis, and
University of Maryland is a separate, distinct institution."
Sonnet 5 is right that those are two separate institutions. It's wrong to fail the claim over it. When I built the labelled set, I read that same comma-separated list as a shortened, recognisable way of naming four affiliations, not as an assertion that one university sits inside another, and graded it supported on exactly that basis. Given room to reason, Sonnet 5 made the stricter call: treat the compression as an error rather than as shorthand.
Four of Sonnet 5's five new errors are this same shape, all against sentences I had graded supported
as fair compression: rz-05, rz-06, sb-04, sb-06. On sb-04, the objection is that a -5.6
point figure is attributed to a harness-model pairing, Codex working with GPT-5.2, not to "one
model" as the sentence has it; I had accepted "one model" as shorthand for one model-harness
configuration. On rz-06, it's that a "Challenger" generates the questions and a "Solver" answers
them, not one single model quietly doing both, as the sentence implies; I had accepted "a model" as a
fair compression of the method. Every one of these objections is a correct reading of the source.
None of them is wrong about the passage. What moves is the tolerance for compression: reasoning made
the models pickier than the benchmark's own construction, not more accurate than it.
That's not a hole in how I labelled the benchmark. unsupported already covers a claim that's
correct in substance but wrong in attribution, scope or wording, which is exactly the shape of these
four objections, and I could have graded any of them that way. I didn't, because I judged the
compression immaterial when I built the set. Reasoning pushed three frontier models to make the
opposite call on the same four sentences, and by the benchmark's own standard they were wrong to.
Whether that stricter standard is the right one for your pipeline is a separate question from whether
it helped on this one.
The other reason not to trust the reasoning arm too much
There's a second problem, and it undercuts the arm's own premise. For GPT-5.4 and Sonnet 5, the verdict call, the second of the two, returns 21 to 23 output tokens on every one of GPT-5.4's 42 rows and 22 to 25 on every one of Sonnet 5's. Not enough visible output to reconsider anything in: it's close to the JSON envelope and the label, nothing else. Whatever those two models decided, the most likely reading is that they decided it in the scratchpad, and the verdict call transcribes it rather than re-deciding on it. That's an inference from output length, not something I can see happening inside the model, and the verdict call isn't free while it does it in JevBench's own accounting: for GPT-5.4 it still averages 35.7% of the reasoning arm's cost and 30.4% of its latency, almost all of it the input side, re-reading the claim and the scratchpad rather than writing a longer answer. Gemini is the exception either way: its verdict call output ranges from 14 to 1453 tokens, so it's doing something closer to actual re-evaluation in that second call.
For two of the three models, calling this "letting them reason before they answer" overstates what happened. It's closer to letting them reason, then asking them to type up what they'd already decided. Whether a genuinely deliberative second call would move the number is a different experiment from the one this run is.
What this changes
The question that mattered was whether the hard-set gap in the first piece was an artefact of forcing the frontier models straight into a schema with no room to think. It has an answer now: no. The gap survives giving them that room, and it survives paying 2.4 to 4.2 times more per claim for it.
It doesn't survive being called significant, on eighteen items and one run each. And the
pre-registered JevBench rule from the first piece still fails here too, though by less than it did
schema-only. That rule says Jev only replaces the check if it's at least as good as the frontier
models on unsupported recall. Jev's recall on unsupported is 78.9%. GPT-5.4's was 89.5%
schema-only and drops to 84.2% with reasoning allowed, so that gap narrows from 10.6 points to 5.3,
even as Jev's hard-18 accuracy lead over GPT-5.4 widens the other way, from 11.1 points to 16.7. Jev
still wins the headline hard-set comparison and still loses the recall rule it was supposed to pass.
Both are true of the same run.
Run this before you trust any before/after shift
If I'd stopped at the headline table, this would read as a clean result: reasoning made two of three models worse. It isn't clean, and the only way to find that out is to run the same check on your own before/after numbers before calling a shift like this real.
Exact McNemar on the discordant pairs only. Count how many items flipped from right to wrong (call it b), how many flipped from wrong to right (call it c), and ignore everything that didn't change. If the two arms were truly equivalent, a flip is a coin flip about which direction it goes. So a lopsided b against c is the signal, and an even split is noise. The p-value is just the two-sided exact binomial on b and c at 0.5.
mcnemar.py in the repo does exactly that, with no dependency beyond the standard library, on any
two JSONL runs over the same items:
python3 mcnemar.py before.jsonl after.jsonl --id-field id --label-field label --predicted-field predicted
Each file is one JSON object per line, matched on the ID field, with a true label and a predicted
label. Those three field names default to id, label and predicted, so if your files already
use those keys you can drop all three flags. Run it before the word "significant" goes anywhere near
a before/after comparison, mine or yours.
What this is not
One run per model per arm, no repeats, no seed, so there's no variance estimate anywhere in this. Test 4 in this series is five runs per model for exactly this reason, and none of today's numbers should be leaned on hard before it runs.
The 200-word cap I put on JevBench's scratchpad is a variable I didn't control for. The schema-only arm had no cap, because it wasn't asked for prose at all. The cap was never breached (the longest scratchpad ran 196 words), but I can't rule out that a longer leash changes anything.
Four of Jev's correct answers on the contested rows sit below its own 0.5 confidence: rz-05 at
0.24, sb-07 at 0.29, sb-06 at 0.37, sb-05 at 0.49. Part of its lead here is Jev being unsure
and right, not confident and right.
Latency ran through OpenRouter for both arms, and this arm makes two round trips, so its routing hop is counted twice against a single-call arm. Jev runs through a different vendor's endpoint entirely, so none of the latency numbers above are a clean cross-vendor comparison, and the two arms ran about 26 hours apart with no serving provider pinned in either.
And the labels are still one person's, checked against the primaries by hand, not by an independent panel. That's the honest caveat on the materiality calls above, too: another builder might draw the line in a different place than I did when I built the set.
The next test in the series asks what happens when the question itself changes: a pinned model
version instead of a moving alias, the compound unsupported label split into two so a stricter
materiality reading and a flat contradiction stop sharing one bucket, and the multi-question fan-out
TypeSafe actually sells instead of one question at a time. Any of the three could move today's
numbers. I wouldn't call the gap settled before they run.
So here's the decision this leaves you with. Don't spend on a reasoning step to fix a reliability problem before you've measured it, on your own failures, with the test above. It cost 2.4 to 4.2 times more here and didn't move the number it was supposed to move. What it did do was make three models pickier than the standard I'd already built the benchmark to. That's worth knowing before you turn reasoning on, not after.
The script that produced every number above is run_llm_reasoning() in run.py, and the analysis
is mcnemar.py and score.py, all at
github.com/TheWayWithin/jev-bench. No API keys needed
to check any of it: the raw runs are already in the repo.
If it saved you an afternoon, buy me a coffee.
Sources
The benchmark: JevBench, reasoning-allowed arm run 21 September 2026, on the same 42 claims as
the first two pieces in this series. openai/gpt-5.4, anthropic/claude-sonnet-5 and
google/gemini-3.1-pro-preview via OpenRouter, one run per model. Jev via TypeSafe, called through
the moving jev-latest alias, not re-run: its figures are the 20 September run. Every number above
comes from those runs, computed by score.py, risk_coverage.py and mcnemar.py, and the
per-claim results, including every model's full reasoning text, are in the repo alongside them.
The labelled set is built from source audits of three papers: SkillsBench (arXiv:2602.12670), ReasoningBank (arXiv:2509.25140) and R-Zero (arXiv:2508.05004).
The first two pieces in this series: Jev is not an LLM, with the accuracy, cost and latency figures the schema-only rows above come from, and GPT-5.4 beats Jev on a confidence threshold, with the calibration analysis this piece doesn't repeat.