Jev 1.13 Tested: Its Router’s Accuracy for 2 Cents per 1,000 Decisions

TypeSafe’s Jev 1.13 only decides; it never writes. Picking AI model sizes for 94 of our tasks, it scored 73 to its own router’s 72, in a fifth of a second and for about 2 cents per 1,000 decisions.

Subject
TypeSafe’s Jev 1.13 tested as an AI model router, against its own router and a free local model
Published
25 SEP 2026
Reading time
9 min

In short

The result
Jev 1.13 picked the right model size for 73 of our 94 tasks. Jev’s router, sold on OpenRouter, got 72: no difference we could detect.
Speed and price
Jev 1.13 answered in a median 0.21 seconds, for about 2 cents per 1,000 decisions. The router took 3.5 seconds and 81 cents.
The free option
Google’s EmbeddingGemma 300M, on our Mac with an add-on trained on our labels, got 60, 13 behind; our strictest test could not confirm that gap.
What to check
Jev 1.13’s surest 70% of answers were 91% right. Its least sure answers are the candidates for a second look, a rule still untested.
In this post · 5 sections
  1. —In short
  2. 01A fast decider in front of slower writers
  3. 02What we tested, and what Jev 1.13 scored
  4. 03Who gets which tasks right
  5. 04Extra context added nothing we could detect; its confidence sorted its answers
  6. 05Where a fast decider fits
The same argument in 3:09. Our graphics, AI narration.

Jev 1.13 (jev-1.13.0), from the start-up TypeSafe AI, is a decision model: it never writes, and only picks an answer from a list, with a probability for each. Asked to pick the right Claude model size for 94 of our tasks, it got 73 right, against 72 for Jev’s router, the version TypeSafe sells on OpenRouter.

A gap of one task is too small for this test to detect. The difference is speed and price: Jev 1.13 answered in 0.21 seconds instead of 3.5, for about 2 cents per 1,000 decisions instead of 81.

A fast decider in front of slower writers

People decide in two ways, fast and slow. Reading the mood of an email takes a second and no effort; filling in a tax form takes an hour of deliberate work. Most AI models work the slow way. They write their answer word by word, and charge for every word.

TypeSafe built Jev to be the fast way for software. It reads a text and a short list of possible answers, and returns a probability for each, charging only for what it reads. That makes it a gatekeeper: a fast decider in front of slow, expensive writers, choosing which of them takes each task.

A task goes to Jev 1.13, which only decides: it picks which of four Claude models should do the work, Haiku, Sonnet, Opus or Fable, in a median 0.21 seconds for about $0.00002 a decision. The Claude models write the answer word by word. Jev is billed for what it reads; its answers are free.FAST DECIDERONE MODEL DECIDES, OTHERS WRITEWHICH MODEL SHOULD DO THIS?DWG Nº 32SCHEMATICTASKa requestFAST: IT ONLY DECIDESwhich model should do this?JEV 1.13picks one, never writes0.21 s a decision$0.00002 a decisionSLOW: THEY WRITE THE WORKCLAUDE HAIKUsmallestCLAUDE SONNETCLAUDE OPUSCLAUDE FABLElargestTHE FAST DECIDERpicks one answer from a list you give itreturns a probability for each answerbilled for what it reads; answers are freeTHE SLOW WRITERSwrite the answer, word by wordreason step by step when neededbilled for what they read and write
A fast decider in front of slow writers: Jev 1.13 never writes, it only picks which Claude model should, in about a fifth of a second. The writing models do the slow, word-by-word work.

TypeSafe’s launch post ties the name to an old observation about coal: as it got cheaper to burn, Britain burned more of it. At about 2 cents per 1,000, a decision is cheap enough to put in front of every task, so how often it is right, and which way it errs, matter more than its price.

What we tested, and what Jev 1.13 scored

In September we tested Jev’s router, listed as typesafe/jev-router on OpenRouter, a marketplace for AI models. A router picks which model handles each request; on 46 tasks it picked the right model size 29 times, and our own keyword rules 19. But it passes each question to another company’s large model at that model’s price, and across 107 calls eight different models answered.

This time we called Jev 1.13 itself on TypeSafe’s own service, and widened the test to 94 tasks. Each task carries a label: the Claude model size our written rules call for, from Haiku, the smallest, through Sonnet and Opus to Fable, the largest. Every label was set before these routers were scored: ours for the older 46 tasks, the majority of three AI raters for the newer 48. The raters did not all agree on 23 of the 94.

We compared three ways to make the call: Jev 1.13; Jev’s router, using its answers from the day before; and Google’s EmbeddingGemma 300M, a small open model we ran free on our Mac, at full precision. It ran plain, and with a trained layer: a small model trained to predict our labels from its output, scored only on tasks it had not seen.

What we measured94 tasks, labelled before scoring

How many times did each pick the right size of model, out of 94?

73Jev 1.13

72Jev’s router

60EmbeddingGemma 300M, trained layer

Jev 1.13 and its router are 1 task apart; three runs of the same questions scored 73, 72 and 76. Each score could be about 9 percentage points higher or lower.

Our tests, 3 and 4 Oct 2026 · 94 tasks from our own work

Right model size out of 94 tasks, median time per decision and price per 1,000 decisions. Jev 1.13 with short rules: 73 right, 0.21 seconds, $0.018. Jev 1.13 shown 8 examples: 76, 0.21 seconds, $0.030. Jev’s router on OpenRouter with short rules: 72, 3.5 seconds, $0.81; shown 8 examples: 78, 2.6 seconds, $1.19. EmbeddingGemma 300M on our Mac with a trained layer: 60, 0.13 seconds, free; untrained: 45, free. Repeat runs of Jev 1.13 scored 73, 72 and 76.THE TRADERIGHT, WAIT AND PRICE, ONE ROW EACH94 OF OUR TASKS, LABELLED FIRSTDWG Nº 334 OCT 2026RIGHT, OF 94model sizeWAITmedian, per decisionPRICEper 1,000 decisionsROUTERJEV 1.13short rules730.21 s$0.018JEV 1.13shown 8 examples760.21 s$0.030JEV’S ROUTERshort rules723.5 s$0.81JEV’S ROUTERshown 8 examples782.6 s$1.19EMBEDDINGGEMMA 300Mtrained layer, our Mac600.13 s$0EMBEDDINGGEMMA 300Muntrained, our Mac450.13 s$0Filled: rented over the internet. Hollow: EmbeddingGemma 300M on our Mac, main processor, warm;0.08 s on the Mac’s graphics chip. Three runs of Jev 1.13: 73, 72, 76. Each score isabout 9 percentage points either way; Jev 1.13 and its router are within that on every setup.
No accuracy gap we could detect, at a fraction of the wait and the price: on 94 routing tasks Jev 1.13 scored 73 to its router’s 72, at 0.21 seconds and about 2 cents per 1,000 decisions against 3.5 seconds and 81 cents. EmbeddingGemma 300M with a trained layer is faster, with no charge per call, and 13 tasks behind.

Shown eight labelled examples, Jev 1.13 and its router scored 76 and 78, again too close to call. The difference is in the wait and the bill. Most of Jev 1.13’s fifth of a second is the trip across the internet: a TypeSafe call that runs no model at all takes 0.15 seconds from our office.

With the model loaded, EmbeddingGemma 300M answers in 0.13 seconds on our Mac’s main processor, or 0.08 on its graphics chip, with no charge per call. Its trained layer got 60 right, 13 behind Jev 1.13. That lead passed our first test but not the stricter correction for running several comparisons at once. Against plain EmbeddingGemma, which got 45, Jev 1.13’s lead survived that correction.

Our AI agents also load packaged skills, reusable instruction sets for particular jobs. Asked which of eight a task needs, if any, Jev 1.13 got all 28 test tasks right, including the 16 that need none.

Who gets which tasks right

A score of 73 and a score of 72 can hide different mistakes, so we compared the routers task by task. Jev 1.13 and its router agree on most tasks: both right on 63, Jev 1.13 alone on 10, the router alone on 9. Against EmbeddingGemma’s trained layer the split is lopsided: 24 tasks only Jev 1.13 got right, 11 only EmbeddingGemma.

The same 94 tasks, two routers at a time. Jev 1.13 and Jev’s router: both right 63, only Jev 1.13 right 10, only the router 9, both wrong 12. Jev 1.13 and EmbeddingGemma 300M with a trained layer: both right 49, only Jev 1.13 24, only EmbeddingGemma 11, both wrong 10. By the size each task needed, Jev 1.13 against EmbeddingGemma: Haiku tasks 26 and 19 of 28, Sonnet 21 and 22 of 33, Opus 19 and 11 of 22, Fable 7 and 8 of 11.TASK BY TASKTHE SAME 94 TASKS, TWO AT A TIMERED: ONLY JEV 1.13 RIGHTDWG Nº 344 OCT 2026Jev 1.13 and Jev’s router6310912Jev 1.13 and EmbeddingGemma 300M, trained layer49241110both rightonly Jev 1.13 rightonly the other rightboth wrongRIGHT, BY LABELLED MODEL SIZEJev 1.13EmbeddingGemmaHAIKU · 28 TASKS2619SONNET · 33 TASKS2122OPUS · 22 TASKS1911FABLE · 11 TASKS78Jev 1.13 leads on Haiku and Opus tasks; on Sonnet and Fable tasks the two are a task apart.Dashed outline: the tasks of that size. 23 of the 94 labels were disputed by our raters.
Jev 1.13 and its router mostly agree, task by task. Against EmbeddingGemma 300M’s trained layer the split is lopsided, 24 tasks to 11, and Jev 1.13’s lead sits in the smallest and the second-largest sizes: trivial edits and hard engineering.

The two fail differently. EmbeddingGemma sends trivial edits, such as a typo fix or a version bump, to a bigger model than they need, and underrates hard engineering, review and debugging. It seems to read a task’s topic rather than its difficulty. Those two sizes hold Jev 1.13’s whole lead: on Haiku tasks it was right 26 times in 28 to EmbeddingGemma’s 19, and on Opus tasks 19 in 22 to 11.

Jev 1.13 has three habits of its own. It sent 8 of 33 Sonnet tasks that sound small, such as adding a loop flag, down to Haiku. It sent fact-checking one size up, reading the word verification in our rubric literally. And it hesitated to pick Fable for strategic calls, which our rubric calls reserved; shown eight examples, it got 3 of those 4 right.

Some of those errors are arguable. Every router did worse on the 23 tasks whose labels our raters disputed, and 10 of Jev 1.13’s 21 errors matched a dissenting rater’s vote. Combining Jev 1.13 with EmbeddingGemma scored 73 to 75, against 73 for Jev 1.13 alone: no gain this test could detect.

Extra context added nothing we could detect; its confidence sorted its answers

We also gave Jev 1.13 more to go on: an automatic rating of each task’s stakes, reversibility, effort, steps and more, added to the question. It scored 75 against 73, inside its own run-to-run noise. On the simpler question of what kind of task each one was, it did worse, 81 against 85.

Two leads cost nothing extra, because they reuse the probabilities Jev 1.13 already returns. The first reads them with a penalty, counting a model one size too small as twice as bad as one size too big. That halved its under-sizing, sending a task to a smaller model than it needs, from 14 tasks to 7. The gain did not clear our pre-set bar.

The second is its confidence. Jev 1.13’s median confidence was 0.87 on answers that turned out right and 0.51 on wrong ones. On the 70% of tasks it was surest about it was right 60 times in 66, or 91%; on its surest half, 44 in 47. The 28 answers it was least sure of held 15 of its 21 mistakes. In September, on a simpler question, the router’s confidence averaged 0.96 when right and 0.94 when wrong.

Keeping only Jev 1.13’s surest answers. All 94 tasks: 78% right. Its surest 90%: 81% right; 80%: 84%; 70%: 91%, 60 of 66; 60%: 93%; 50%: 94%, 44 of 47; 40%: 95%; 30%: 96%. Its median confidence was 0.87 on right answers and 0.51 on wrong ones. In our September test, on the task-kind question, Jev’s router averaged 0.96 on right answers and 0.94 on wrong ones.CONFIDENCEKEEP ONLY ITS SUREST ANSWERSHOW MANY ARE RIGHT? · 94 TASKSDWG Nº 35OBSERVED, NOT PRE-SET0%50%100%RIGHT78%ALL81%90%84%80%91%70%93%60%94%50%95%40%96%30%share of the 94 tasks kept, surest first60 of 66HOW SURE IT SAID IT WASJev 1.13 · model size · median 0.87 when right, 0.51 when wrong · 94 tasksJev’s router · Sept, task kind · mean 0.96 when right, 0.94 when wrong · 44 scoredWhich answers to send for a second look is a rule still to be tested on fresh tasks.
Confidence that sorts: on the 70% of tasks Jev 1.13 was surest about it was right 60 times in 66. In September, on a simpler question, its router’s scores barely differed between right and wrong answers.

That makes Jev 1.13’s least sure answers the natural candidates for a second look, by a person or a bigger model; whether that second look improves the result is still untested.

For engineersThe method, for anyone repeating it: runs on 3 and 4 October 2026, plans registered before any call was scored.

Jev 1.13: TypeSafe’s systemone endpoint, model jev-1.13.0, one pick-one question per task with the four tier descriptions as options. Jev’s router: the typesafe/jev-router listing on OpenRouter, answers from 3 October. EmbeddingGemma 300M: Google’s embeddinggemma-300m, bf16 weights converted to full-precision sentence-transformers format, 768-dimension output, on our Mac, not the 4-bit build our site search runs. Its trained layer is a logistic-regression probe on frozen vectors, five-fold cross-validated.

Exact McNemar tests, Holm correction across secondary tests. Primary, Jev 1.13 against the router on short rules: 73 against 72, p = 1.0, difference +1.1 points, 95% interval −8.5 to +10.6. Against the trained layer: 73 against 60, p = 0.041, Holm 0.29. Against untrained EmbeddingGemma: 73 against 45, Holm p = 0.0002. Scope as extra context: 75 against 73, p = 0.80. Penalty reading: weighted error score 27 against 37, p = 0.20, Holm 1.0. Keeping the surest answers was observed, not pre-registered. Paid spend for both rounds: $0.043, at $0.042 per million input tokens, output free.

Where a fast decider fits

The jobs that suit Jev 1.13 share one shape: the same small choice from a fixed list, made many times, where a fifth of a second and 2 cents per 1,000 matter, and the unsure cases can be set aside. Writing falls outside that shape by design.

Jev 1.13 · what we measured, and what TypeSafe says

Good fit

  • Choosing the right size of model for a task: 73 of 94 on ours, with its router’s accuracy at a fraction of its wait and price.
  • Picking from a short, fixed menu, such as which packaged skill a task needs: 28 of 28.
  • Ranking its own answers by confidence: its surest 70% were 91% right.
  • TypeSafe’s own examples: sorting support tickets by urgency, flagging claims that need an adjuster (use-case map, confidence routing).

Poor fit

  • Writing anything. It doesn’t generate text; in TypeSafe’s words, “there are other models for that” (known weak spots).
  • Sums, counting or several steps of reasoning. TypeSafe says to keep the maths in ordinary code.
  • Tasks whose difficulty hides behind modest wording: it sent 8 of 33 Sonnet tasks to Haiku.
  • Strategic calls that need the largest model, unless shown examples: it hesitated on 4 of 11.

Whether to rent that judgement or run a free model yourself is the subject of our post on local AI on a Mac. It keeps search on the machine and rents the judgement calls.

A fast decider earns its place when it is cheap, quick, and its confidence tracks its mistakes. On our tasks Jev 1.13 matched a router that costs 45 times as much per decision, and its least sure answers held most of its errors. Measure both on your own work before it starts deciding for you.

AGNIZAR
Production AI · System architecture · Fractional CTO

Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.

Book an AI Architecture Review