Skip to main content

Jev, a cheap AI, did two thirds of my filing: same result as Claude, 63% cheaper

Jamie Watters

Operational resilience and AI delivery practitioner.

Published: 25 September 2026•10 min read
#jev#typesafe#claude#second-brain#evaluation
A card reading: A cheap AI did two thirds of my filing. Same result as Claude, 63% cheaper. Jev passing unsure notes to Claude 86.7% at $0.0034 a note; Claude alone 86.7% at $0.0091 a note; keyword search with no AI 76.0%. But it kept Claude's mistakes.
Tests 1, 1b and 1c, run 23 and 24 September 2026 on notes from my own vault.

Who this is for: you pay for AI to sort something over and over (emails, tickets, notes, documents) and you'd like the bill smaller without the results getting worse. You'll come away with what that saved on my notes, the catch, and a test to run on your own work.

Skip it if you've never paid for an AI model by the call. Most of this won't matter to you yet.

On the length: the result and the catch are in the first three paragraphs. The rest is how sure you can be of them, what didn't help, and the test.


A cheap AI model can do most of a sorting job and hand the hard cases to an expensive one. I tested that on 75 of my own notes. Jev, a model TypeSafe released on 15 September, filed each note into a folder. Where it was confident, its answer stood. Where it wasn't, the note went to Claude Sonnet 5, which costs about 76 times as much per note. Jev handled about two thirds itself.

The result was the same as giving every note to Claude: 65 of the 75 filed where I'd filed them. The model bill was 63% lower. (Every figure here is from my Test 1c results, listed under Sources.)

The catch is that it kept Claude's mistakes too. On the six notes where Jev was confident and wrong, Claude usually picked the same wrong folder, so passing notes up couldn't have fixed them.


Three tests, one question each

My notes live in a vault, filed into seven folders. Five are projects, paid work, article drafts, things I've read and reference. The other two are my book and Efformism, a philosophy framework I'm developing. (The tests also offered an eighth answer, delete.) Right now a Claude session proposes where each note goes and I say yes or no. Jev looked made for the job: you give it text and a fixed set of answers, and it returns one answer and a confidence score. Claude sessions built and ran the tests; I approved each set of rules before it ran.

Test 1, on 23 September, gave Jev and Claude each note and a one-line description of each folder. Jev lost: on their most common answers over five runs, 65.1% to 77.1% (12 notes to 2, p = 0.013).

Test 1b, on 24 September, added ten examples to every note. They were the most similar notes already filed, found by a keyword search, each shown with its folder. Jev rose to 83.1%. But the search on its own, taking the folder most of the ten were in, got 74.7%, and on 83 notes the test couldn't tell the two apart.

Test 1c asked whether Jev's hand-off works. It scored 75 notes that no earlier test had scored, and the rules were written down before any of them were run. One caveat belongs here rather than at the end: 54 of those 75 had already appeared as examples, with their folders, in earlier runs, including the rewording runs described below. The models keep nothing between calls, so that doesn't affect the Jev or Claude results. It does weaken the rewording result.


The hand-off

Jev answers every note. If its confidence is 0.80 or more, its answer stands; below that, Claude's answer is used. Jev's confidence is a separate score from the probability it gives its top answer, and here it was a little lower on 289 of 375 answers. Claude was asked to state its own confidence. Claude answered every note so the two could be compared, and each Jev run was paired with one Claude run. In use Claude would only see the notes Jev passes on, so the cost is Jev on every note plus Claude on those.

75 notes Right, most common answer over 5 runs Share of answers passed on Model charges per note
Keyword search only 76.0% (1 run) n/a $0
Jev 85.3% n/a $0.000119
Jev, passing unsure notes to Claude 86.7% 32.5% $0.00340
Claude alone 86.7% n/a $0.00909

Bar chart of how often each setup filed one of 75 notes where I had filed it, using the most common answer over five runs (one run for the search, which always gives the same answer): keyword search only 76.0% at no cost; Jev 85.3% at $0.00012 a note; Jev passing unsure notes to Claude 86.7% at $0.0034 a note; Claude alone 86.7% at $0.0091 a note.

Jev alone got 64 notes right. Passing its unsure notes to Claude fixed 3 and broke 2 that Jev had right: 65. On the notes Jev kept, the two had the same majority answers, right ones and wrong ones. On the notes Jev passed on, the pair's answer simply was Claude's. So the two ended on the same 65.

That match is the finding, and I should be plain about how much it proves. The rule I approved was that a standard statistical test must not show the hand-off worse than Claude, at no more than half the cost. It passed. But on 75 notes that rule could not have failed by much: a hand-off 7 notes worse than Claude would still have passed it, because I didn't set how much worse was acceptable. The evidence is the observed result, not the rule. With no disagreements in 75, the true disagreement rate could still be up to about 4% (a one-sided 95% bound, worked out afterwards). Averaged over the five runs rather than by most common answer, the pair got 85.9% and Claude 86.4%.


Confident isn't the same as right

Bar chart. On 75 notes, the share of answers given at or above the confidence line that were wrong, against a 5% limit set before the test: Jev 11.9%, 30 of 253; Jev with a yes or no question per folder 9.3%, 15 of 161; Claude, for comparison, 5.4%, 14 of 259.

Of the answers Jev kept, 30 of 253 were wrong: 11.9%, against a 5% limit I set before the test. On Test 1b's notes it had been 5.5%. Jev never changed its folder for a note between runs, so the 30 are six notes, each wrong in every run. For comparison only, since the rule applied to Jev, Claude's confident answers were wrong 14 times in 259, 5.4%.

Those six notes decide the verdict. Two are my career fact sheet and my contacts register, filed under paid work; both models said reference. Two are content ideas: a review of a source, which both put with things I've read, and an article brief, which both put under projects. The Claude session that wrote up the results thinks several look like my filing habits rather than model mistakes. That's my call, and I haven't made it. If four of the six turn out to be labels I'd change, Jev's rate falls to 4.0% and passes. Until then no labels change, because relabelling after a model disagrees is exactly the temptation the rule exists to stop.

So by the rule I set, Jev doesn't file on its own, even with me as the backstop. It would act on two thirds of its answers. Of all its answers, 8.0% would have been acted on wrongly (Test 1c verdict).


What the test couldn't separate

Narrower questions. TypeSafe recommends breaking a broad question into narrow ones and combining the answers in code. So one setup asked Jev eight yes-or-no questions in one call, one per answer. The rule for combining them, set in advance: the highest yes wins, but a note is passed on unless the top yes is at least 0.80 and no other reaches 0.50. It filed 85.3%, the same as the single question, one note better and one worse. It passed on 57.1% of its answers, almost all because no yes reached 0.80.

Rewording. Jev isn't fine-tuned on customer data; you shape it through what you send with each request. So the Claude session running the tests reworded Jev's folder descriptions from my filing rules, working only on Test 1b's 83 older notes. There Jev went from 69 right to 74. On the 75 test notes it went from 64 to 66, p = 0.5. An independent check found that some of the rewording went beyond my written rules and some was prompted by particular older notes, and the 54 overlapping notes were among the examples in the rewording runs. So I don't lean on that result.

Jev against the search. Jev got 64 notes right and the search 57. Jev was right on 13 that the search got wrong, and wrong on 6 that it got right (p = 0.167). That's a lead the test can't confirm, not a tie. The search gave no confidence, so it couldn't pass anything on; a search that hands off when its vote is close wasn't tested.


What this test can't show

Most of my notes were filed by a Claude session and accepted by me, so the labels may lean Claude's way, which makes agreement with Claude partly built in. 54 of the 75 test notes had appeared as examples in earlier runs, including the rewording runs. None of them was scored until the test itself, but they weren't new to it either. The sample is balanced across folders on purpose, so it isn't the mix my real inbox gets. The test notes include none from my book or Efformism folders, because Test 1 used all four of each. The hand-off was to Claude without a reasoning step. And the Test 1c runs took six minutes on one day, so they say nothing about whether either model behaves the same next month.


Set it up this way

BEFORE YOU LET AN AI FILE THINGS FOR YOU

1. Pick filed items to tune on and a separate set to
   test on. I used 83 and 75, up to 15 per folder,
   drawn with a fixed seed. Keep the test set out of
   every tuning step, including as examples.
2. Strip anything added at filing time (tags, labels).
3. For each item, find its 10 most similar filed items,
   never the item itself. I used BM25 (k1 1.5, b 0.75)
   over title plus first 1,500 characters, and showed
   each example as title, first 500 characters, folder.
4. Give the model the item, a line describing each
   folder, and the 10 examples. Get a folder and a
   confidence. Jev returns its own; I asked Claude to
   state one. Run it five times.
5. Before any test run, freeze the question wording
   and write down three numbers: the confidence line
   (I used 0.80), the most wrong answers you'll accept
   above it (I used 5%), and how much worse than the
   strong model the hand-off may be. I didn't set the
   third, and should have.
6. At or above the line the answer stands. Below it,
   pass the item to the strong model or to you.
   Cost = cheap model on every item
        + strong model on the items passed on.
7. Score once, on the test set: the hand-off against
   the strong model alone, and the cheap model alone
   against the search's vote.
8. Count wrong answers at or above the line. Over your
   limit? Someone still checks those. If a label is
   wrong, fix it for next time; don't fix it to pass.

For now Jev doesn't file my notes on its own. Passing its unsure third to Claude gave the same result as Claude alone for 63% less, and kept Claude's mistakes along with its answers.

Sources

The notes are private, so the raw material lives in my vault at tools/jev-stack-tests/. The rules for each test were written before its runs (test1-preregistration.md, test1b-preregistration.md, test1c-preregistration.md). Every Test 1c figure is in results/test1c-summary.json, which analyse1c.py regenerates from the raw runs, or in the verdict results/2026-09-24-test1c-verdict.md, with its deviations and the two independent checks. Test 1b: results/2026-09-24-test1b-verdict.md. Test 1: results/2026-09-23-test1-verdict.md. The code and per-note results with every note's text, title and path removed are being added to the public jev-bench repository. Jev jev-1.13.0 via TypeSafe; anthropic/claude-sonnet-5 via OpenRouter. Model charges: Jev from TypeSafe's published input price, $0.042 per million tokens with output free (TypeSafe's model docs, read 20 and 24 September 2026); Claude from OpenRouter's reported charge. Retrieval ran locally and isn't costed.

Why examples: in the public BANKING77 experiment, a single run, Jev scored 92.4% on 77 kinds of banking request with 24 retrieved labelled examples per message. Its README, which it says GPT-6 Astra Light wrote, adds: "we did not measure a retrieval-only classifier, so we cannot quantify Jev's incremental contribution over that alternative."

The earlier pieces this builds on: Test an AI model before you rely on it, the six mistakes and the checklist these tests follow; and I said Jev beat GPT-5.4. Five reruns can't tell them apart, the fact-checking test where Jev's confident wrong answers first turned up.

No paywall, no sponsors. If this saved you some time, you can buy me a coffee.

☕ Buy me a coffee

Get the next one in your inbox

I build with AI in the open and write up what held and what didn't. Real numbers, the failures before the wins.

Share this post