Skip to main content

Jev is a narrow tool with four clear uses. The $10m framing is not one of them.

Jamie Watters

Operational resilience and AI delivery practitioner.

Published: 24 September 202610 min read
#jev#decision-models#small-business#ai-hype
A card quoting the episode: $10 million businesses, $100 million businesses. Beside it, the same episode's guest saying to use it for routing only. The expensive queue it claims to unlock costs $1.29 a month on the old model for a thousand enquiries, so nothing was rationed. A narrow tool with four clear uses.
Quotes from the episode's auto-transcript. Costs from my own 42-claim benchmark, 20 September 2026.

Who this is for. You run a small business or work for yourself, you keep seeing videos telling you some new AI tool has unlocked a business you could start, and you would like to know whether that is true before you spend a weekend on it. By the end you will have the one sum that tests any of those claims, and the four conditions under which this particular kind of tool is actually worth using.

Skip it if you want a guide to using Jev, or a technical explanation of how it works. Neither is here.

On the length. About ten minutes. If you only want the method, skip to "The check, before you build anything": it is five numbered steps you can run in an afternoon.


A podcast episode about a new AI model opens with the host briefing his guest on what he wants from the next half hour. The auto-transcript of the episode has him asking for this, and the figures in it are his, not mine (https://youtu.be/4mTLpuQpB80):

"three or four insane use cases so that people can walk away from this episode with like productivity, making money, just like, you know, even boring use cases that could become, you know, $10 million businesses, $100 million businesses" (auto-transcript, https://youtu.be/4mTLpuQpB80)

The model is Jev, from TypeSafe. It is genuinely interesting and I have spent a week benchmarking it. So I did the thing the format does not invite, which is put each claim next to a number.

Two of the three headline claims do not survive. The third one the guest dismantles himself, on camera, and the packaging carries on regardless.


Claim one: these are businesses that are now unlocked

The episode's framing is "startup ideas that are now unlocked", and the mechanism is stated plainly later on:

"How do I find a business with an expensive queue of incoming information and then just put Jev at the front of that queue?"

Expensive queue. That is the whole economic argument, and it is checkable.

In my own 42-claim benchmark on 20 September, Jev came out at $0.000025 per check against GPT-5.4's $0.001289. Raw runs: https://github.com/TheWayWithin/jev-bench

So take a plumber getting a thousand enquiries a month, which is far more than a sole trader gets. At GPT-5.4's measured $0.001289 per check from that same table (https://github.com/TheWayWithin/jev-bench), reading and sorting every single one of them costs $1.29 a month.

That is the expensive queue. A pound and a bit.

No sole trader was rationing at that price. Nobody looked at a pound a month and decided the business was uneconomic. Whatever has stopped these ideas being built, it was never the bill, so the price falling cannot be what unlocks them.


Claim two: the ideas themselves are new

They are not, and this is the gentler correction. The three the episode lands on are faster replies to trade enquiries, finding the good thirty seconds in an hour of footage, and sorting a contact form by which enquiry is real.

All three were buildable last year, with the expensive model, at a cost nobody would have noticed. I say that as someone who has just done the arithmetic rather than as a matter of opinion.

What actually stops them is not in the episode at all:

  • The enquiries are not in one place. They arrive by WhatsApp, by phone, through a website form and two lead marketplaces. Something has to gather them before anything can sort them, and that something is the build.
  • This model writes no text. It picks from a list. So the "reply to the customer" step, the one the whole idea depends on, needs a template or a person. The model cannot do it.
  • The labels are the job. Deciding what "a real project with a real budget" means, precisely enough for a machine to apply it, is most of the work. My own benchmark's published weaknesses include exactly that: it answers the question you wrote rather than the one you meant.

Claim three: the guest already corrected this one

Here is the part that makes the episode better than most, and also makes the packaging worse.

Asked for more use cases, the guest volunteers a failure. He had wired the model up to produce buy, hold or sell signals on Bitcoin, once a minute:

"it does not seem to be doing well, which shows that this model is great, but it does have some regressions. I would not put this model in front of like your stock portfolio or Bitcoin or anything like that. This is just for like routing or other sort of decisions like that where it doesn't need insane and model intelligence."

And:

"it's not the best in all of the situations"

That is an accurate, useful, unprompted description of what the tool is for. Most people demonstrating something new do not volunteer the case where it fell over.

It also flatly contradicts the framing it sits inside, quoted at the top of this piece from the same transcript (https://youtu.be/4mTLpuQpB80). The guest is saying: narrow tool, routing only, keep it away from anything where being wrong costs you. The episode's opening is saying: insane use cases, hundred million dollar businesses, stick around to the end.

Both things are in the same half hour. Only one of them is in the title.


So what is it actually for

You do not have a conversation with it. You hand it something (an email, a transcript, a form submission) and the answers it is allowed to give. It picks one, or scores it, or returns a probability, and tells you how sure it is. It writes no text and shows no reasoning.

A two-column comparison. A chatbot takes a question in words and returns text it wrote, and is good at judgement and explanation. A decision model takes an item plus the list of answers it may choose from, returns one of those answers and how sure it is, and is good at sorting the same question thousands of times quickly.

Think of a traffic officer rather than an adviser. It tells you which lane, thousands of times an hour, and it is wrong occasionally in ways you can measure.

That narrowness is why it is fast and cheap. It is also why it is the wrong tool for anything needing judgement. That is not a criticism, it is the design.


Where it genuinely makes sense

This is the section the episode owes its audience and does not deliver. The guest gave the rule in one line: routing, not judgement. Here is what that means in practice.

It fits when all four of these are true:

  1. The same decision, over and over. Thousands of items, one question each. Not one hard call a week.
  2. The answers are a fixed list you can write down in advance. Spam or not. Urgent, normal, or ignore. Inside the service area or outside it.
  3. Being wrong is cheap to correct. A misfiled email costs you a click. A wrongly quoted job costs you the job.
  4. Something else does the acting. The model picks the lane. A template, a rule or a person still has to send the reply, book the slot or quote the price.

It does not fit when:

  • The decision needs information that is not in front of it, which is what killed the Bitcoin test.
  • The answer needs to be written rather than chosen.
  • The cost of a confident mistake lands on a customer.
  • You cannot say, before you start, exactly what the right answer looks like.

That is a narrow tool with clear uses, which is exactly what the guest said it was. It is not a business in a box, and nothing about it makes money by existing.


What the price drop does change

There is a real effect here and it deserves separating from the sales pitch, because dismissing it would be the opposite error.

Cheapness bites where the volume is enormous and the decisions are per-event: scoring every message in a live stream, asking twenty questions of one item instead of one, a browser agent deciding something every few hundred milliseconds. At that scale the old price does hurt and the new one does not.

That is a much narrower claim than the one the episode opens with. It is also true, and it is the honest version of the pitch.


The check, before you build anything

This costs an afternoon and saves a quarter. It works for this model or any other.

1. Collect a hundred real examples. From your own inbox, forms or footage. Not invented ones. Invented examples are clean and real ones are not, which is the whole reason the test works.

2. Write down the right answer for each, before you run anything. You already know them.

3. Decide your pass mark now, in writing, before you see a single result. Two numbers: the share it must get right, and the share of confidently-wrong answers you can live with. Write them down before you look. It is the step that stops you moving the goalposts once you have seen the answer.

4. Run them through and count three things: how many it got right, what it cost, and separately, how many it got wrong while saying it was sure. That third number is the one that reaches a customer.

5. Work out what the alternative costs you. Five hundred enquiries a month at four minutes each is about thirty-three hours. Put your own hourly rate against that. It is the number the automation has to beat, and it is almost never the API bill.

Four steps. One, collect a hundred real examples from your own inbox, forms or footage, not invented ones. Two, write down the right answer for each. Three, run them through and count how many are right and at what cost. Four, count the confidently wrong ones separately, because those are the ones that reach a customer.

If it is wrong in ways that cost you nothing to correct, you have a business. If it is wrong in ways that lose you a client, you have a demo.


Why I trust that test more than I trust my own headline

Because it has already taken one off me.

On 20 September I published that this model beat GPT-5.4 by eleven points at checking whether a source supports a claim. I reran it five times. The gap came out at 3.3 points, which on eighteen items is 0.6 of a single claim, and not significant.

The same runs say something sharper than either headline, and the table is in the same repo (https://github.com/TheWayWithin/jev-bench). On the clean examples I constructed myself, the frontier models scored 100% and Jev 91.7%. On the eighteen real published sentences, the frontier models fell to 66.7% and Jev held at 77.8%. The expensive models look perfect on invented material and come apart on real material. This one is the other way round.

That inversion is the actual finding, and I had it backwards in the first version of this article, which is its own small lesson in how easily a number gets retold into its opposite.

Every number is open at https://github.com/TheWayWithin/jev-bench, including the five reruns that took my own headline away from me.

The tool is real and the ideas are worth having. What is not worth having is the sentence "this unlocks new businesses" from anyone who has not multiplied the old price by your actual volume.

No paywall, no sponsors. If this saved you some time, you can buy me a coffee.

☕ Buy me a coffee

Get the next one in your inbox

I build with AI in the open and write up what held and what didn't. Real numbers, the failures before the wins.

Share this post