I said Jev beat GPT-5.4. Five reruns can't tell them apart, at a fiftieth of the cost.
Jamie Watters
Operational resilience and AI delivery practitioner. Technology since 1985.

Who this is for. You're wondering whether Jev, or any cheap new AI model, could do your checking instead of GPT-5.4. Or you've picked an AI tool because it won a test you ran once. By the end you'll know what Jev actually did against GPT-5.4 on my hard checks, and you'll have a check you can run before trusting any AI comparison, mine included.
Skip it if you only want the verdict. On this test I can't tell Jev and GPT-5.4 apart, and Jev costs a fiftieth as much.
On the length. About ten minutes. The result is in the first screen. The rest is how I know, and the check.
On 20 September I published that Jev, TypeSafe's new model, beat GPT-5.4 by 11.1 points on the hard half of my fact-checking benchmark, at a fiftieth of the cost. Two pieces later I published that the lead grew to 16.7 points when the big models were allowed to think first.
Both numbers came from running each model once.
So on 22 September I ran every model five times on the same 42 claims. Jev doesn't beat GPT-5.4. On the hard claims its lead averages 3.3 points, and that isn't statistically significant against GPT-5.4 or either of the other two frontier models. On this test I can't tell them apart. The single run I built two articles on caught Jev at its best.
That isn't the same as saying they're equal. Eighteen claims is too few to show that, and one of them is probably labelled wrong. Take that one claim out and Jev and GPT-5.4 average exactly the same.
Two more things are worth knowing if you're weighing one against the other. Jev still costs a fiftieth as much. And it lets more wrong claims through: of the 19 wrong claims in the set, it caught 15 and GPT-5.4 caught 18.
This is me correcting my own headline, in public, with every raw run open so you can check I'm not making the same mistake twice. And if you've ever picked any AI tool because it won a test you ran once, this is what that decision can be standing on.
What I ran
Same benchmark as before, JevBench: 42 claims, each paired with the source passage it cites, and the job is to say whether the passage supports the claim. 18 are hard: real sentences from one of my published articles, labelled from audits of the papers they cite. Nine are wrong, six are right, and three cite a passage that doesn't cover them. The wrong ones are mostly not flat contradictions. They take a fact from the wrong version of a paper, quote something the paper doesn't say, or claim more than it shows. The other 24 are easy: controls I built on purpose, where the right answer is plain.
Four systems. Jev, from TypeSafe, and three frontier models: GPT-5.4, Claude Sonnet 5 and Gemini 3.1 Pro. Each one answered all 42 claims five times, with exactly the prompt I used in the first article. No temperature or seed was set on either side, also as in the first article.
I wrote down in advance how I'd score it and what would count as the lead holding, and committed that file to the JevBench repo 30 seconds before the first call. The runs took 19 minutes, all in one window, so they don't measure drift from one day to the next. The whole thing, both tests in this piece, cost $2.25.
The numbers
| all 42, mean (range) | the easy 24 | the hard 18, mean (range) | hard 18, the single run I published | |
|---|---|---|---|---|
| Jev | 83.8% (83.3–85.7) | 91.7% every run | 73.3% (72.2–77.8) | 77.8% |
| GPT-5.4 | 87.1% (85.7–88.1) | 100% every run | 70.0% (66.7–72.2) | 66.7% |
| Sonnet 5 | 85.7% (83.3–88.1) | 100% every run | 66.7% (61.1–72.2) | 66.7% |
| Gemini 3.1 Pro | 84.3% (81.0–85.7) | 100% every run | 63.3% (55.6–66.7) | 55.6% |
Across the JevBench runs, Jev still has the highest average on the hard 18. But look at the ranges. Jev's worst run, 72.2%, is the same as the best run of GPT-5.4 and of Sonnet 5. Gemini never got above 66.7%. Jev scored 77.8% in one of its five runs and 72.2% in the other four.
I also ran the pinned version of Jev five times, and it turned out to be the same model. Counting those JevBench runs too, 77.8% came up in four runs out of ten, and the ten-run average is 74.4%. Either way, the run I published was at the top of Jev's range.
On averages, Jev leads GPT-5.4 by 3.3 points, Sonnet 5 by 6.7 and Gemini by 10.0. One claim on the hard set is worth 5.6 points, so the lead over GPT-5.4 is 0.6 of one claim.

I tested whether any of those gaps is real with an exact McNemar test, which only looks at the claims where one system was right and the other wrong. A claim counts as right if a system got it right in at least three of its five runs. Jev against GPT-5.4: Jev alone right on 3 claims, GPT-5.4 alone right on 2, p = 1.000. Against Sonnet 5, p = 0.500. Against Gemini, p = 0.250.
None of those shows a real difference. They can't show the two are equal either. A rough 95% interval on the JevBench gap between Jev and GPT-5.4 runs from about 19 points worse to 30 points better. On 18 claims, that's how little the test can see.
The rule I wrote down beforehand calls this result "shrinks": Jev's average still leads, but the run ranges overlap. So the sentence "Jev beats the frontier models on the hard claims" can't be written on this evidence. I wrote it on 20 September.
Why one run was so far off
Of the 18 hard claims, Jev gives the same answer on all five runs for 16 of them. All of its spread
comes from two claims, sb-06 and sb-07, each right three times in five. The frontier models
wobble the same way, on one to three claims each.
On a set this small, that's the whole story. One claim flipping moves the score 5.6 points. The first article's 11.1-point lead was two claims: Jev 14 of 18, GPT-5.4 12. In that JevBench run Jev got one claim more than it does in four runs out of five. GPT-5.4 got one fewer than it does in three runs out of five. Those two claims are the 11.1 points.
One thing didn't move at all. The three frontier models scored 100% on the easy 24 in all fifteen of their JevBench runs. Jev scored 91.7% in all ten, missing the same two controls every time. That part of the first article holds exactly: whatever edge Jev has is on the hard claims, and nowhere else.
This also changes what the third piece, on letting the big models think first, can claim. Its 16.7-point lead was measured from Jev's 77.8% JevBench run, which is now known to be Jev's best. The reasoning runs weren't repeated. Its other finding, that room to think didn't improve any of the frontier models, is also one run per model, so it gets the same caution as everything else here.
Did I ask Jev the wrong way?
This piece was meant to be about a different mistake. The plan was to find out whether I'd been asking Jev badly and then blaming the model. So I tested three changes to the question, one at a time, five runs each.
Pinning the version. I'd been calling jev-latest, a name that moves when a new release ships.
Every jev-latest call on 22 September was answered by jev-1.13.0 anyway, so calling
jev-1.13.0 directly was calling the same model. It made no detectable difference: its averages sat
inside the spread of the other five runs, and its majority answers were identical.
Splitting a vague label. My "unsupported" answer covers two different things: a claim the source contradicts, and a claim that's roughly right but altered. I split it into two options. Jev scored 1.4 points lower. The three frontier models, given the same split once each, each fixed two claims and broke two others.
Asking several narrow questions at once. This is the shape TypeSafe recommends: several small questions about the same passage in one call, billed once. It cost the same, $0.000026 per claim against $0.000025. It scored 6.2 points lower on all 42. That's the biggest change in either test, and it still isn't significant: p = 0.219.

So none of the three changes did better than the plain question. That's three changes, though, not every way of asking Jev.
The split did turn up one thing worth saying. Claim sb-06 says part of a paper's reported gain "is
unattributed". The paper doesn't use that word. It says the gain could partly come from giving the
model more text, and asks for the controls that would settle it. I labelled the claim supported, as
a fair paraphrase. Given an "altered" option, every system chose it for that claim, every time. On
that claim the likeliest thing that's wrong is my label.
It matters more than one claim should. Jev got sb-06 right three times in five and GPT-5.4 never
did, so it's one of the claims Jev's lead rests on. Take it out of the JevBench results and both
average 74.1% on the other 17.
Jev doesn't give the same answer twice
Asked the identical question five times, Jev changed its confidence on 28 of the 42 claims and its answer on 2. Every answer that changed had a confidence below 0.47 on every run.
That confidence number needs a word of explanation. It isn't the probability of the answer Jev
picked. On sb-06, in one run, Jev gave its answer a probability of 0.53 and reported a confidence
of 0.28. It's a separate score, and here it's the one that tracked whether an answer would stay put.
That matters if you store an answer and reuse it rather than asking again. On this evidence it's safe at 0.5 confidence and above: no claim that reached 0.5 on any run changed its answer, across 420 calls under both version names. That's for the plain single question. With several narrow questions, two answers flipped with a score sitting right around 0.5. And it's a rule about answers staying the same, not about them being right.
The five-run check
This is what I'd do before believing any "A beats B" result on an AI tool, a model or a prompt, including my own. On JevBench it took 19 minutes.
THE FIVE-RUN CHECK
1. Run each option five times on the same inputs. Same prompt, same settings.
2. Write down every run's score, as well as the average.
3. If the winner's worst run doesn't beat the loser's best run,
you don't have a winner yet. (A quick sanity check, not a test.)
4. Work out what one item is worth. On 18 items, one is 5.6 points.
If the gap is one or two items, say so in those words.
5. Count the items where one option was right and the other wrong.
Fewer than six, and no test can call the gap significant,
however lopsided it is.
6. Only then run an exact McNemar test on those items.
7. A gap that isn't significant is not a tie. It means you can't tell.
Step 3 is the rule I wrote down before these runs. Jev failed it. For step 6, mcnemar.py in the
repo compares two runs, one answer per item, and needs nothing beyond standard Python:
python3 mcnemar.py before.jsonl after.jsonl
Each file is one JSON object per line, with an id, the true label and the predicted answer. To
compare five runs against five, run it on each pair of runs, or count a claim as right when it's
right in three of five, the way tests34.py does for these runs.
Five runs tell you how much a result moves on your items. They don't tell you how it would go on different items. That needs more items, ideally written by someone else.
So should you use Jev instead of GPT-5.4?
Here's what five runs support. On the real published sentences, I can't tell Jev and GPT-5.4 apart. On the easy controls Jev is worse: the frontier models were perfect in every run and Jev missed the same two every time. And the price gap is enormous. Jev cost $0.000025 per claim on 22 September, against $0.001289 for GPT-5.4 and $0.004940 for Gemini, about a fiftieth to a two-hundredth of the price per call on this task.
But look at which mistakes each one makes. Of the 19 wrong claims in the JevBench set, counting a claim as caught if a model caught it in at least three of five runs, Jev caught 15 and GPT-5.4 caught 18. Sonnet 5 and Gemini caught 16 each. Jev's edge on the hard claims came from the other direction: GPT-5.4 failing three claims that were actually right. That difference isn't significant either (p = 0.25), but it points the wrong way for a checker, whose main job is catching what's wrong.
So if a wrong claim slipping through is expensive for you, Jev isn't a drop-in replacement for GPT-5.4 on this evidence. If you need a very cheap first look and something else checks what matters, it's worth testing on your own work. Either way, treat a Jev answer below 0.5 as one that might change if you ask again.
The strongest case against all of this: 18 hard claims, five runs, all from my own writing, labelled by me. The next test in this series asks other people for their published mistakes for exactly that reason.
The first and third pieces in this series now carry a note pointing here.
Every number above comes from the raw runs at
github.com/TheWayWithin/jev-bench.
python3 tests34.py analyse reproduces the accuracy, cost and significance figures without an API
key, and the rest can be worked out from its per-claim table and the raw files. The file I committed
before the first call is there too, so you can check I didn't move the goalposts after seeing the
results.
A two-claim lead from one run was 0.6 of a claim across five. That's what one run was worth.
If it saved you an afternoon, buy me a coffee.
Sources
The benchmark: JevBench, Tests 3 and 4, run 22 September 2026 from 18:29 to 18:48 EDT. 38 runs,
1,596 calls, no failed calls. Jev via TypeSafe as jev-latest and jev-1.13.0; openai/gpt-5.4,
anthropic/claude-sonnet-5 and google/gemini-3.1-pro-preview via OpenRouter. Scoring rules
committed in results/2026-09-22-tests-3-4-preregistration.md before the first call. Full results
and limits: results/2026-09-23-tests-3-4-verdict.md.
The labelled set is built from source audits of three papers: SkillsBench (arXiv:2602.12670), ReasoningBank (arXiv:2509.25140) and R-Zero (arXiv:2508.05004).
The earlier pieces in this series. First, Jev is not an LLM, the single-run comparison this piece corrects. Second, GPT-5.4 beats Jev on a confidence threshold, the calibration analysis. Third, Paying your AI to think longer doesn't fix its mistakes, the reasoning test. Its 16.7-point lead was measured from the same single Jev run.
This piece was reviewed before its final version by six outside AI models, reading the live draft against the repo. They caught two factual errors, a Gemini figure and the order of the series, and all six objected to my first headline word, "tie".