Back to Research
Paper
Test an AI model before you rely on it: six mistakes I made testing Jev
22 min read
One run said Jev beat GPT-5.4 by 11.1 points. Five runs could not tell them apart. How to test an AI model before you rely on it, from the six mistakes behind that headline.
Read the paperThe series, in reading order
- Jev is not an LLM. On the hard half it beat GPT-5.4, Gemini 3.1 Pro and Sonnet 5, at a fiftieth of the cost.18 min read
- GPT-5.4 beats Jev on a confidence threshold, once tested at its own values.26 min read
- Paying your AI to think longer didn't fix its mistakes. It cost up to 4.2x more.14 min read
- I said Jev beat GPT-5.4. Five reruns can't tell them apart, at a fiftieth of the cost.13 min read
- Jev, a cheap AI, did two thirds of my filing: same result as Claude, 63% cheaper12 min read