Test an AI model before you rely on it: six mistakes I made testing Jev
Jamie Watters
Operational resilience and AI delivery practitioner. Technology since 1985.

Who this is for. You're choosing an AI tool, model or prompt for work that matters, and someone has shown you a comparison where one beats another. It might be the vendor's, a colleague's, or one you ran yourself. By the end you'll have a checklist you can run on that comparison in an afternoon, and you'll know which of its results to believe.
Skip it if you only want to know whether Jev can replace GPT-5.4. Short answer: on my hardest fact-checks I can't tell them apart, Jev costs about a fiftieth as much, and it lets more wrong claims through. The full re-test is here.
On the length. About eighteen minutes. The checklist is near the end, under "The checklist". The rest is where each line of it came from.
On 20 September I published that Jev, a new AI model from TypeSafe, beat GPT-5.4 by 11.1 points on the hardest fact-checks I could give it, at a fiftieth of the cost. The result came from my own benchmark, JevBench.
It didn't hold. On 22 September I ran every model five times on the same claims. On JevBench's hard claims Jev's lead over GPT-5.4 averaged 3.3 points, which is 0.6 of one claim, and no test I could run could tell the two apart. Jev also let more wrong claims through: of the 19 wrong claims in the set, it caught 15 and GPT-5.4 caught 18.
The models didn't change between those two dates. My testing did.
Everything that went wrong is something anyone can do when they compare two AI tools. So this is the whole arc, from headline to correction, with every raw run open, including the runs that proved me wrong. Six mistakes, what each one did to the result, and the checklist I'd now run before relying on any AI comparison, mine included.
What I was testing
Jev isn't a chatbot. It doesn't write anything. You give it a piece of text and a question with a fixed set of answers, and it hands back one of those answers and a probability. TypeSafe charges $0.042 per million input tokens for it, with output free (TypeSafe's model docs).
The job I gave it is one I actually need done. Before anything publishes under my name, a script checks it. That script can tell whether a number has a source attached. It can't tell whether the source says the number. So the question for each claim was: here is a sentence, here is the passage it cites, does the passage support the sentence as written? Three possible answers: supported, not supported, or the passage doesn't cover it.
JevBench has 42 claims. 18 are hard: real sentences from an article of mine that had to be rewritten, each paired with the passage it cites. 24 are easy: controls I built, where one detail is changed and the right answer is plain. Four systems answered all of them: Jev, and three frontier models, GPT-5.4, Claude Sonnet 5 and Gemini 3.1 Pro.
Here's how the headline moved, piece by piece, all on the same JevBench claims:
| run | what I published about the hard 18 | what it turned out to be |
|---|---|---|
| 20 September, one run each | Jev beats every frontier model by 11.1 points | Jev's best run, against GPT-5.4's worst |
| 21 September, frontier models allowed to think first, one run each | Jev's lead over GPT-5.4 is 16.7 points | measured from that same best run |
| 22 September, five runs each | Jev's lead averages 3.3 points, not significant | the test can't tell them apart |
Six outside AI reviewers read the correction against the raw files before its final version. Most of what follows is what they, and the five reruns, made me see.
Mistake 1: I ran it once
AI models don't always give the same answer when you ask the same question twice. I knew that, and I tested as though I didn't.
On the 18 hard JevBench claims, Jev scored 77.8% in one of its five runs and 72.2% in the other four. The run I published was the 77.8%. GPT-5.4's five runs sat between 66.7% and 72.2%, and the run I published scored 66.7%, its lowest. Put Jev's best beside GPT-5.4's worst and you get 11.1 points. Put the averages side by side and you get 3.3.
I also ran a pinned version of Jev five times, and it turned out to be the same model. Counting those JevBench runs too, 77.8% came up in four runs out of ten, and the ten-run average is 74.4%. Either way, the run I published was at the top of Jev's range.

On a small test, a gap is counted in items, not percentages. With 18 claims, each one is worth 5.6 points. My 11.1-point lead was two claims. In that one JevBench run, Jev got one more claim right than it usually does, and GPT-5.4 got one fewer. That's the whole headline.
What to do instead. Run each option at least five times and write down every score, not just the average. If the winner's worst run doesn't beat the loser's best run, you don't have a winner yet. And spread the runs over more than one day: my five-run JevBench test ran in one 19-minute window, so it can't show whether a model behaves differently on another day.
Mistake 2: I headlined a score that wasn't the job
Before the first run, I wrote a pass mark into the JevBench README. Jev would replace the check in my publishing gate only if it caught at least 80% of the wrong claims, and at least as many as the frontier models. I wrote it down in advance so the result couldn't become a story about whichever way it landed.
On day one, in the first JevBench run, Jev caught 15 of the 19 wrong claims, 78.9%. GPT-5.4 caught 17. By my own rule the answer was a first filter, not a replacement. That answer was in the body of the article. The headline used a different number, accuracy on the hard 18, where Jev was ahead.
Five runs made the gap between those two numbers plain. Counting a claim as caught when a model caught it in at least three of its five JevBench runs, Jev caught 15 of the 19 wrong claims, GPT-5.4 18, and Sonnet 5 and Gemini 16 each. Jev's lead on the hard claims had come from the other direction: GPT-5.4 failing three claims that were actually right.

Neither gap is significant on this few claims. But for a checker the two kinds of error don't cost the same. If it wrongly flags a good sentence, I look at it, find it's fine, and lose five minutes. If it wrongly passes a bad one, the bad sentence goes out under my name. An accuracy figure adds those two together and hides the expensive half inside the cheap half.
What to do instead. Before running anything, write down the job and pick the score that matches it. For a checker, that's the share of wrong answers it catches, with false alarms reported beside it. For something that drafts text, it might be the share of drafts you'd send without editing. Headline that score. Put overall accuracy underneath, never on top.
Mistake 3: I tested it on my own writing, labelled by me
All 18 hard JevBench claims are sentences from one article of mine. The passages come from the papers that article cited. The right answers were set by one person, me, from audits of those papers made inside my own publishing process, and checked by hand against the papers.
That has two costs, and the five runs showed me both.
The first is that one of my labels is probably wrong. Claim sb-06 says part of a paper's reported
gain "is unattributed". The paper doesn't use that word. I'd labelled the sentence supported, as a fair
paraphrase. When I gave the models an extra answer for "roughly right but altered", every system chose
it for that claim, on every run. And it isn't any claim: Jev got sb-06 right three times in five and
GPT-5.4 never did, so it's one of the claims Jev's lead rested on. Take it out and, on the other 17
JevBench claims, Jev and GPT-5.4 both average 74.1%.
The second is quieter. My 18 hard sentences are mistakes that had already got past me: past my writing, my automated check and a first round of review. So they're a sample of what I miss, not a sample of what a checker sees in normal work. That's the right set for asking whether a tool catches my blind spots. It's the wrong set for predicting how it does on anyone else's.
What to do instead. Test on items someone else wrote, or on real failures from your own work that you didn't pick for the test. Have a second person label them without seeing your answers, and report how often the two of you agreed. If you can't agree on the right answer, no model can be scored against it.
Mistake 4: I published before testing the setup
My first run made three setup choices without testing any of them. I called Jev by a name that moves when a new version ships, and didn't record which version answered. I asked it one question per claim, when TypeSafe's own guidance recommends several narrow questions about the same text in one call. And I made the frontier models answer straight away, with no room to reason first, which every reviewer of that piece raised.
Every one of those could have moved the result, and I published the result anyway. When I tested them afterwards, two made no detectable difference and the third scored lower. None of the three gaps is statistically significant.
Letting GPT-5.4, Sonnet 5 and Gemini think before answering improved none of them on the hard 18, at 2.4 to 4.2 times the cost per claim. That was one run per model, and none of the changes is significant, so the honest reading of that JevBench run is "didn't help", not "made it worse".
TypeSafe recommends asking several narrow questions at once. On JevBench that cost the same as one question: $0.000026 per claim against $0.000025. It scored 6.2 points lower on all 42. That isn't significant either (p = 0.219), and it's one way of splitting the question, not every way.

Pinning the version made no detectable difference, because every call on 22 September was answered by the same version anyway. But my first run didn't record which version answered it, so I can't prove what I tested on 20 September. The evidence says it was very probably the same model. The file can't show it.
What to do instead. Treat every recommended setup as a guess until you've tested it against the plain one, whether it comes from the vendor's docs, a colleague, or the common advice to "turn on reasoning". Pin the model version, write down every setting, and record which version actually answered. A vendor's advice may be right for their examples. Whether it's right for your job is something only your test can tell you.
Mistake 5: I misread the confidence score, twice
Both Jev and the frontier models return a confidence number with each answer. It's tempting to treat that number as permission: trust the confident answers, check the rest. I got it wrong in both directions.
The first time, I tested GPT-5.4's confidence at six round thresholds, 0.60 up to 0.95, chosen before I'd looked at its answers. All 42 of its JevBench answers sat between 0.94 and 0.99. Five of my six thresholds were below every one of them, so the results came out flat, and I published that its confidence told you nothing. It does. Tested at the values it actually returns, keeping only its 0.99 answers gave 64.3% coverage at a 3.7% error rate. That correction is its own piece.
The second time was about Jev. Asked the identical question five times, Jev changed its answer on 2 of the 42 JevBench claims, and every answer that changed had a confidence below 0.47 on every run. Across 420 calls, no claim that reached 0.5 on any run ever changed its answer. My first draft of the correction presented 0.5 as a line above which you could trust Jev. Four of the six outside reviewers said I'd turned a rule about answers staying the same into a rule about answers being right.
They were right, and the data shows why. Claim sb-03 is a quotation in my article that is really a
section heading joined to a sentence from further down the paper. On JevBench's plain question, Jev
called it supported in all eleven of its runs, at a confidence of 0.98 or 0.99 every time. It's
wrong. All four models got it wrong in every plain-question run. Sonnet 5 and Gemini only got it
right when I offered a separate answer for "altered". A confident answer that stays the same can
still be wrong.
What to do instead. Test a confidence score at the values the model actually returns, not at round numbers you picked. Then test it against whether answers were right, not just whether they stayed the same. If it passes both, you can use it to decide what to check first. It still doesn't tell you what you can stop checking.
Mistake 6: I called "can't tell" a tie
The correction first went live with the word "tie" in its title. All six outside reviewers objected, and they were right.
Here's what the five JevBench runs support. Counting a claim as right when a model got it right in at least three of five runs, Jev was right on 3 hard claims where GPT-5.4 was wrong, and GPT-5.4 on 2 where Jev was wrong. An exact McNemar test, which looks only at those disagreements, gives p = 1.000. A rough 95% interval on the gap (a Wald interval, the simple normal approximation) runs from about 19 points worse for Jev to 30 points better.
That doesn't say the two are equal. It says that on 18 claims the test can't rule out a gap of about 20 points either way. "Tie" claims something the test can't measure. The word I'd written down before the runs for this outcome was "shrinks": Jev's average still leads, but the ranges overlap. I should have used it.
What to do instead. When a test finds no significant gap, say "can't tell them apart", and say how big a gap it could have missed. If you need to show two options are as good as each other, decide before you run how close counts as "as good", and use enough items to show it. As a rough guide from this data: if two models disagree on the same share of items as Jev and GPT-5.4 did, pinning the gap to within 5 points either way takes something like 430 items, not 18.
What a fair test needs
Put the six fixes together and you get the shape of a test worth relying on. None of it is advanced. Most of it is writing things down before you look.
Before you run: the job, the score that matches it, the pass mark and what "as good as" means, all written down and dated. Items you didn't write, labelled by two people who didn't compare notes. Enough items that one item isn't worth five points.
While you run: at least five runs of each option, on more than one day, with every version and setting recorded, and the recommended setup tested against the plain one.
When you read it: every run's score, not just the average. The disagreements counted before any significance test. Errors broken down by the kind that costs you most. And "can't tell" written as "can't tell".
JevBench, as it stands, does about half of that. Five runs, a pass mark set in advance, the score that matches the job reported, the setup tested, and exact tests on the disagreements. It doesn't have outside items, a second labeller, a margin for "as good as", or runs on different days. So it can tell you that my first headline was wrong. It can't yet tell you whether Jev is as good as GPT-5.4 at this job.
The checklist
This is what I'd now run before relying on any "A beats B" result for an AI tool, model or prompt. That includes a vendor's result, a colleague's, and mine. Copy it.
BEFORE YOU RELY ON "A BEATS B"
Before you run anything
1. Write down the job, and the score that matches it.
For a checker: of the wrong answers, how many did it catch?
Overall accuracy goes beside that score, never above it.
2. Write down the pass mark, and how close counts as "as good as".
Date it. Don't change it after you've seen results.
3. Use items you didn't write, or real failures you didn't pick.
Have a second person label them without seeing your labels.
Write down how often you agreed.
4. Work out what one item is worth: 100 divided by the number of items.
If one item is worth more than a couple of points, get more items.
5. Pin the model version. Write down every setting.
While you run
6. Run each option at least five times, on more than one day.
Record which version actually answered.
7. Test the recommended setup (the vendor's, or "turn on reasoning")
against the plain one. Don't assume it's better.
When you read the results
8. Write down every run's score, not just the average.
If the winner's worst run doesn't beat the loser's best, no winner yet.
(A quick sanity check, not a test.)
9. Count the items where one option was right and the other wrong.
Fewer than six, and an exact McNemar test can't call the gap
significant, however lopsided it is.
10. Run an exact McNemar test on those items. Reduce each option's runs
to one answer per item first (right if right in at least three of
five), or compare one run against one run.
11. Comparing against more than one option? Name the main comparison
before you run. Treat the others as a first look, not a result.
12. Not significant means "can't tell them apart". Never "tie".
Only the margin you set in step 2, and enough items, can show
"as good as".
13. Break the errors down by kind. Look hardest at the ones that cost most.
14. If there's a confidence score, test it at the values it actually
returns, and against whether answers were right, not just whether
they stayed the same.
Step 9 is arithmetic, not opinion. With five disagreements all going one way, the exact test's p-value is 0.0625, which misses the usual 0.05 line. With six, it's 0.031. If your comparison has fewer than six disagreements, no exact test on those items will make the gap significant.
For step 10, mcnemar.py in the JevBench repo compares
two files, one answer per item each, with nothing beyond standard Python:
python3 mcnemar.py before.jsonl after.jsonl
Each file is one JSON object per line, with an id, the right answer as label, and the model's
answer as predicted. To use five runs, write one file per option holding each item's majority
answer first; tests34.py in the same repo does that reduction for the JevBench runs.
Where this checklist stops
The strongest objection comes from people who evaluate models for a living, and it's fair. Five runs and an exact test on 18 items is still a weak design. A proper analysis would use every run for every item rather than a majority vote, test for equivalence against a margin set in advance, and run on hundreds of items. It would also deal with running several comparisons at once. I compared Jev with three frontier models and tested each gap on its own, and the more gaps you test, the likelier one of them looks real by chance. None of mine was significant, so it didn't bite here, but a future "Jev wins" would need one comparison named in advance. By that standard, even a fully corrected JevBench is a pilot, not a verdict.
I agree with all of it. The checklist isn't a substitute for that work. It's the floor underneath it: the checks that would have stopped my headline. They're cheap enough to run on any comparison you're shown, and a comparison that fails them isn't worth the more careful analysis until it has more items.
The other fair objection is that I've built a checklist from one small benchmark on one kind of task. That's true too. Every mistake above happened to me on fact-checking. Whether they're the mistakes that matter most for, say, choosing a coding assistant is something I can't show from this data.
Send me your published mistakes
The fix for mistake 3 is claims I didn't write. The next version of JevBench will be built from other people's published mistakes: sentences that went out, cited a source, and turned out not to say what the source said, then were corrected.
If you have one, I'd like it. What's useful is the published sentence, the source it cites, the passage, what was wrong with it, and a link to the correction. Open an issue on github.com/TheWayWithin/jev-bench with those five things. The new test will be designed and written down before anything runs, and a result that goes against Jev gets written up the same week and with the same prominence as one that goes for it.
One run said 11.1 points. Five runs said 0.6 of one claim. The models were the same both times.
If it saved you an afternoon, buy me a coffee.
Sources
The benchmark: JevBench, at github.com/TheWayWithin/jev-bench.
First run 20 September 2026, one run per model. Reasoning-allowed run 21 September, one run per
frontier model. Five runs per system on 22 September, 18:29 to 18:48 EDT, scoring rules committed
before the first call. Jev via TypeSafe; openai/gpt-5.4, anthropic/claude-sonnet-5 and
google/gemini-3.1-pro-preview via OpenRouter. Verdicts: results/2026-09-22-test2-verdict.md and
results/2026-09-23-tests-3-4-verdict.md. python3 tests34.py analyse reproduces the five-run
figures without an API key.
The labelled set is built from source audits of three papers: SkillsBench (arXiv:2602.12670), ReasoningBank (arXiv:2509.25140) and R-Zero (arXiv:2508.05004).
Jev's price: TypeSafe's model docs, read 20 September 2026.
The four pieces this one draws on, in the order they were published: Jev is not an LLM, the single-run comparison; GPT-5.4 beats Jev on a confidence threshold, the confidence correction; Paying your AI to think longer didn't fix its mistakes, the reasoning test; and I said Jev beat GPT-5.4. Five reruns can't tell them apart, the five-run correction.