Is a model built for decisions better at them than general-purpose models?
Jev: Twenty Times Faster, Seven Points Behind
Oct 10, 2026
Faster and cheaper, yes. More accurate, no. TypeSafe sells its decision model, Jev, as calibrated ("higher confidence means higher accuracy"), and we found no published measurement of that claim. We tested it on 799 yes/no decisions with known answers, under a plan registered before the run, beside two general-purpose models. Jev answered in about a fifth of a second, around 20 times faster than DeepSeek V4 Flash and 9 times faster than GPT-6.1 Sol, and it cost the least per answer. It was also less accurate: 6.6 points behind DeepSeek V4 Flash and 8.2 behind GPT-6.1 Sol on the same items. Whether its confidence matches its accuracy, the claim we registered, landed too close to our pre-set line to call either way.
Each panel has its own scale. Response time is the median of 3,297 answers per model, observed rather than registered, and includes the network. Accuracy is on the headline items. The numbers are in the table below. Jev (TypeSafe) DeepSeek V4 Flash GPT-6.1 SolA calibrated model sits on the dashed line. Each point is a confidence band (50–60%, 60–70% … 90–100%); bands with fewer than 10 answers are left out, which is why the general-purpose models have fewer points: they almost always said 90% or more. Hover a point for its count.
What stands out
Calibration, the registered question: inconclusive. Jev's calibration error was 0.045, and its 95% interval, 0.031 to 0.068, straddles our pre-set line of 0.05. Under the rule we registered, that verdict is "inconclusive": the result neither supports the claim nor rules it out.
Accuracy, the second registered question: lower, and the gap is clear. On the same items, Jev was 6.6 points less accurate than DeepSeek V4 Flash (interval 4.7 to 8.6) and 8.2 points less accurate than GPT-6.1 Sol (6.2 to 10.2).
Speed and cost, observed: Jev's median answer took 0.18 seconds, against 3.6 for DeepSeek V4 Flash and 1.6 for GPT-6.1 Sol, and its slowest answers were barely slower than its typical ones. It also cost the least: about 3 cents per 1,000 answers, against 15 cents and $1.18.
Put together: choosing Jev over DeepSeek V4 Flash saved about 12 cents per 1,000 decisions and got about 66 more of those 1,000 wrong.
Jev did better than guessing. Always answering "yes" would have scored 67.7% on these items; Jev scored 86.9%.
When Jev was unsure, it said so. Where it reported 70–80% confidence, it was right 85% of the time, so it leaned underconfident rather than over. DeepSeek V4 Flash went the other way: on its 46 answers at 80–90% confidence, it was right 41% of the time.
On items written to be genuinely ambiguous (never scored), Jev stayed near the middle: 79% of its answers fell between 30% and 70%. The two general-purpose models answered firmly instead.
At TypeSafe's documented hold level (act only at 70% confidence or above), Jev acted on 75% of the headline items and was right on 94.8% of those. DeepSeek V4 Flash reached the same accuracy, 94.7%, while answering all of them.
On the weak spots TypeSafe documents, reported separately: Jev got arithmetic questions right half the time (90 answers), the same as a coin flip, while both comparison models were above 95%.
How it was built
Every question is a yes/no decision with an objective answer fixed before any model saw it. No answer key comes from another model's judgement.
Two kinds of item count toward the headline. One kind asks about real AI coding-agent runs from our Harness Battery, each already checked against the workspace and read by hand. The other kind asks whether a shell command, in a described workspace, would permanently destroy data, with the answer worked out by a deterministic simulator.
A third set aims at weak spots that TypeSafe's own documentation lists. It is reported separately and never counted in the headline.
Jev was called through TypeSafe's own API, in its documented yes/no format. It returns one number: the probability of yes. We read its confidence as the distance of that number from 50%. That reading is ours; Jev returns no separate confidence.
Two general-purpose models answered the same items with a stated probability: DeepSeek V4 Flash and OpenAI's GPT-6.1 Sol. They are comparisons, not the subject of the test.
Every item was answered three times by each model: 9,891 answers in all.
The test, the pass line and the wording of each possible verdict were registered on the Open Science Framework before the run (osf.io/qg6nc), and the analysis was run once.
The pass line for calibration is ours, because TypeSafe publishes none: an expected calibration error of at most 0.05.
The numbers
Measure (headline items)
Jev
DeepSeek V4 Flash
GPT-6.1 Sol
Answers scored
2,377
2,282
2,397
Accuracy
86.9%
94.7%
95.1%
Accuracy, 95% interval
84.5–89.1%
93.0–96.0%
93.5–96.7%
Calibration error (lower is better)
0.045
0.035
0.021
Calibration error, 95% interval
0.031–0.068
0.023–0.051
0.009–0.036
Brier score (lower is better)
0.096
0.046
0.035
Gave no usable answer
0
115
0
Answered exactly 50%, so not scored for accuracy
20
0
0
Median response time (observed, not registered)
0.18 s
3.6 s
1.6 s
Slowest 10% of answers took over (observed)
0.22 s
15.4 s
3.8 s
Measured cost per 1,000 answers (described only)
$0.03
$0.15
$1.18
What it does not show
"Inconclusive" is a result, not a delay. Jev's calibration error sits close to the line, so we do not expect a larger rerun to settle it. A new test would need its own registration.
The 0.05 line is ours. TypeSafe publishes no calibration target, so a different reasonable line could give a different verdict.
Our reading of Jev's confidence is our own, because Jev returns only a probability.
DeepSeek V4 Flash gave no usable answer on 115 headline items. Those are left out of its scores, not counted as wrong, so its accuracy is measured on slightly fewer items.
These are our items, about coding-agent work and destructive commands. They are not TypeSafe's own workflows, and we did not test its cost or speed claims.
Speed was not part of the registered test, and we said we would not test TypeSafe's speed claims. We recorded every response time anyway and report them as observations. They include the network: Jev was called directly, and the two general-purpose models through OpenRouter, which adds some delay.
Cost is described, not tested: it is the usage each API reported, priced at the rates on 9 October 2026. Jev's output is free, so its cost depends only on the size of the question.
The detailed findings beyond the two registered questions (the confidence bands, the ambiguous items, the weak-spot set and the hold level) were not part of the registered test and are reported as observations.
Run 9–10 October 2026 against jev-1.13.0, the version Jev's API reported on every answer. We report what the models did; we say nothing about intent.
Results as of Oct 10, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: