Jev 1.13 Tested: Its Router’s Accuracy for 2 Cents per 1,000 Decisions
TypeSafe’s Jev 1.13 only decides; it never writes. Picking AI model sizes for 94 of our tasks, it scored 73 to its own router’s 72, in a fifth of a second and for about 2 cents per 1,000 decisions.
- Subject
- TypeSafe’s Jev 1.13 tested as an AI model router, against its own router and a free local model
- Published
- 25 SEP 2026
- Reading time
- 9 min
- Film
- Watch · 3:09
In short
- The result
- Jev 1.13 picked the right model size for 73 of our 94 tasks. Jev’s router, sold on OpenRouter, got 72: no difference we could detect.
- Speed and price
- Jev 1.13 answered in a median 0.21 seconds, for about 2 cents per 1,000 decisions. The router took 3.5 seconds and 81 cents.
- The free option
- Google’s EmbeddingGemma 300M, on our Mac with an add-on trained on our labels, got 60, 13 behind; our strictest test could not confirm that gap.
- What to check
- Jev 1.13’s surest 70% of answers were 91% right. Its least sure answers are the candidates for a second look, a rule still untested.
In this post · 5 sections
Jev 1.13 (jev-1.13.0), from the start-up TypeSafe AI, is a decision model: it never writes, and only picks an answer from a list, with a probability for each. Asked to pick the right Claude model size for 94 of our tasks, it got 73 right, against 72 for Jev’s router, the version TypeSafe sells on OpenRouter.
A gap of one task is too small for this test to detect. The difference is speed and price: Jev 1.13 answered in 0.21 seconds instead of 3.5, for about 2 cents per 1,000 decisions instead of 81.
A fast decider in front of slower writers
People decide in two ways, fast and slow. Reading the mood of an email takes a second and no effort; filling in a tax form takes an hour of deliberate work. Most AI models work the slow way. They write their answer word by word, and charge for every word.
TypeSafe built Jev to be the fast way for software. It reads a text and a short list of possible answers, and returns a probability for each, charging only for what it reads. That makes it a gatekeeper: a fast decider in front of slow, expensive writers, choosing which of them takes each task.
TypeSafe’s launch post ties the name to an old observation about coal: as it got cheaper to burn, Britain burned more of it. At about 2 cents per 1,000, a decision is cheap enough to put in front of every task, so how often it is right, and which way it errs, matter more than its price.
What we tested, and what Jev 1.13 scored
In September we tested Jev’s router, listed as typesafe/jev-router on OpenRouter, a marketplace for AI models. A router picks which model handles each request; on 46 tasks it picked the right model size 29 times, and our own keyword rules 19. But it passes each question to another company’s large model at that model’s price, and across 107 calls eight different models answered.
This time we called Jev 1.13 itself on TypeSafe’s own service, and widened the test to 94 tasks. Each task carries a label: the Claude model size our written rules call for, from Haiku, the smallest, through Sonnet and Opus to Fable, the largest. Every label was set before these routers were scored: ours for the older 46 tasks, the majority of three AI raters for the newer 48. The raters did not all agree on 23 of the 94.
We compared three ways to make the call: Jev 1.13; Jev’s router, using its answers from the day before; and Google’s EmbeddingGemma 300M, a small open model we ran free on our Mac, at full precision. It ran plain, and with a trained layer: a small model trained to predict our labels from its output, scored only on tasks it had not seen.
What we measured94 tasks, labelled before scoring
How many times did each pick the right size of model, out of 94?
73Jev 1.13
72Jev’s router
60EmbeddingGemma 300M, trained layer
Jev 1.13 and its router are 1 task apart; three runs of the same questions scored 73, 72 and 76. Each score could be about 9 percentage points higher or lower.
Our tests, 3 and 4 Oct 2026 · 94 tasks from our own work
Shown eight labelled examples, Jev 1.13 and its router scored 76 and 78, again too close to call. The difference is in the wait and the bill. Most of Jev 1.13’s fifth of a second is the trip across the internet: a TypeSafe call that runs no model at all takes 0.15 seconds from our office.
With the model loaded, EmbeddingGemma 300M answers in 0.13 seconds on our Mac’s main processor, or 0.08 on its graphics chip, with no charge per call. Its trained layer got 60 right, 13 behind Jev 1.13. That lead passed our first test but not the stricter correction for running several comparisons at once. Against plain EmbeddingGemma, which got 45, Jev 1.13’s lead survived that correction.
Our AI agents also load packaged skills, reusable instruction sets for particular jobs. Asked which of eight a task needs, if any, Jev 1.13 got all 28 test tasks right, including the 16 that need none.
Who gets which tasks right
A score of 73 and a score of 72 can hide different mistakes, so we compared the routers task by task. Jev 1.13 and its router agree on most tasks: both right on 63, Jev 1.13 alone on 10, the router alone on 9. Against EmbeddingGemma’s trained layer the split is lopsided: 24 tasks only Jev 1.13 got right, 11 only EmbeddingGemma.
The two fail differently. EmbeddingGemma sends trivial edits, such as a typo fix or a version bump, to a bigger model than they need, and underrates hard engineering, review and debugging. It seems to read a task’s topic rather than its difficulty. Those two sizes hold Jev 1.13’s whole lead: on Haiku tasks it was right 26 times in 28 to EmbeddingGemma’s 19, and on Opus tasks 19 in 22 to 11.
Jev 1.13 has three habits of its own. It sent 8 of 33 Sonnet tasks that sound small, such as adding a loop flag, down to Haiku. It sent fact-checking one size up, reading the word verification in our rubric literally. And it hesitated to pick Fable for strategic calls, which our rubric calls reserved; shown eight examples, it got 3 of those 4 right.
Some of those errors are arguable. Every router did worse on the 23 tasks whose labels our raters disputed, and 10 of Jev 1.13’s 21 errors matched a dissenting rater’s vote. Combining Jev 1.13 with EmbeddingGemma scored 73 to 75, against 73 for Jev 1.13 alone: no gain this test could detect.
Extra context added nothing we could detect; its confidence sorted its answers
We also gave Jev 1.13 more to go on: an automatic rating of each task’s stakes, reversibility, effort, steps and more, added to the question. It scored 75 against 73, inside its own run-to-run noise. On the simpler question of what kind of task each one was, it did worse, 81 against 85.
Two leads cost nothing extra, because they reuse the probabilities Jev 1.13 already returns. The first reads them with a penalty, counting a model one size too small as twice as bad as one size too big. That halved its under-sizing, sending a task to a smaller model than it needs, from 14 tasks to 7. The gain did not clear our pre-set bar.
The second is its confidence. Jev 1.13’s median confidence was 0.87 on answers that turned out right and 0.51 on wrong ones. On the 70% of tasks it was surest about it was right 60 times in 66, or 91%; on its surest half, 44 in 47. The 28 answers it was least sure of held 15 of its 21 mistakes. In September, on a simpler question, the router’s confidence averaged 0.96 when right and 0.94 when wrong.
That makes Jev 1.13’s least sure answers the natural candidates for a second look, by a person or a bigger model; whether that second look improves the result is still untested.
For engineersThe method, for anyone repeating it: runs on 3 and 4 October 2026, plans registered before any call was scored.
Jev 1.13: TypeSafe’s systemone endpoint, model jev-1.13.0, one pick-one question per task with the four tier descriptions as options. Jev’s router: the typesafe/jev-router listing on OpenRouter, answers from 3 October. EmbeddingGemma 300M: Google’s embeddinggemma-300m, bf16 weights converted to full-precision sentence-transformers format, 768-dimension output, on our Mac, not the 4-bit build our site search runs. Its trained layer is a logistic-regression probe on frozen vectors, five-fold cross-validated.
Exact McNemar tests, Holm correction across secondary tests. Primary, Jev 1.13 against the router on short rules: 73 against 72, p = 1.0, difference +1.1 points, 95% interval −8.5 to +10.6. Against the trained layer: 73 against 60, p = 0.041, Holm 0.29. Against untrained EmbeddingGemma: 73 against 45, Holm p = 0.0002. Scope as extra context: 75 against 73, p = 0.80. Penalty reading: weighted error score 27 against 37, p = 0.20, Holm 1.0. Keeping the surest answers was observed, not pre-registered. Paid spend for both rounds: $0.043, at $0.042 per million input tokens, output free.
Where a fast decider fits
The jobs that suit Jev 1.13 share one shape: the same small choice from a fixed list, made many times, where a fifth of a second and 2 cents per 1,000 matter, and the unsure cases can be set aside. Writing falls outside that shape by design.
Jev 1.13 · what we measured, and what TypeSafe says
Good fit
- Choosing the right size of model for a task: 73 of 94 on ours, with its router’s accuracy at a fraction of its wait and price.
- Picking from a short, fixed menu, such as which packaged skill a task needs: 28 of 28.
- Ranking its own answers by confidence: its surest 70% were 91% right.
- TypeSafe’s own examples: sorting support tickets by urgency, flagging claims that need an adjuster (use-case map, confidence routing).
Poor fit
- Writing anything. It doesn’t generate text; in TypeSafe’s words, “there are other models for that” (known weak spots).
- Sums, counting or several steps of reasoning. TypeSafe says to keep the maths in ordinary code.
- Tasks whose difficulty hides behind modest wording: it sent 8 of 33 Sonnet tasks to Haiku.
- Strategic calls that need the largest model, unless shown examples: it hesitated on 4 of 11.
Whether to rent that judgement or run a free model yourself is the subject of our post on local AI on a Mac. It keeps search on the machine and rents the judgement calls.
A fast decider earns its place when it is cheap, quick, and its confidence tracks its mistakes. On our tasks Jev 1.13 matched a router that costs 45 times as much per decision, and its least sure answers held most of its errors. Measure both on your own work before it starts deciding for you.
Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.