Is a model built for decisions better at them than general-purpose models?

Jev: Twenty Times Faster, Seven Points Behind

Oct 10, 2026

Faster and cheaper, yes. More accurate, no. TypeSafe sells its decision model, Jev, as calibrated ("higher confidence means higher accuracy"), and we found no published measurement of that claim. We tested it on 799 yes/no decisions with known answers, under a plan registered before the run, beside two general-purpose models. Jev answered in about a fifth of a second, around 20 times faster than DeepSeek V4 Flash and 9 times faster than GPT-6.1 Sol, and it cost the least per answer. It was also less accurate: 6.6 points behind DeepSeek V4 Flash and 8.2 behind GPT-6.1 Sol on the same items. Whether its confidence matches its accuracy, the claim we registered, landed too close to our pre-set line to call either way.

Full paper: SILT-RP-009 · DOI 10.5281/zenodo.23287930 — the complete method, every table and the registration.

Speed and accuracy on the same 799 questionsJev (TypeSafe)DeepSeek V4 FlashGPT-6.1 SolMedian response time (shorter is faster)Jev (TypeSafe) · Median response time (shorter is faster): 0.18 s0.18 sDeepSeek V4 Flash · Median response time (shorter is faster): 3.6 s3.6 sGPT-6.1 Sol · Median response time (shorter is faster): 1.6 s1.6 sAccuracyJev (TypeSafe) · Accuracy: 86.9%86.9%DeepSeek V4 Flash · Accuracy: 94.7%94.7%GPT-6.1 Sol · Accuracy: 95.1%95.1%
Each panel has its own scale. Response time is the median of 3,297 answers per model, observed rather than registered, and includes the network. Accuracy is on the headline items. The numbers are in the table below.
Stated confidence against share correct, headline items50%60%70%80%90%100%30%40%50%60%70%80%90%100%How sure the model said it wasHow often it was rightdashed line = perfectly calibratedbelow it = overconfidentJev (TypeSafe): said 56% sure on average, right 55% of the time (233 answers)Jev (TypeSafe): said 65% sure on average, right 69% of the time (371 answers)Jev (TypeSafe): said 75% sure on average, right 85% of the time (332 answers)Jev (TypeSafe): said 85% sure on average, right 91% of the time (403 answers)Jev (TypeSafe): said 96% sure on average, right 99% of the time (1,038 answers)Jev (TypeSafe)DeepSeek V4 Flash: said 82% sure on average, right 41% of the time (46 answers)DeepSeek V4 Flash: said 99% sure on average, right 96% of the time (2,228 answers)DeepSeek V4 FlashGPT-6.1 Sol: said 64% sure on average, right 49% of the time (39 answers)GPT-6.1 Sol: said 74% sure on average, right 64% of the time (64 answers)GPT-6.1 Sol: said 82% sure on average, right 69% of the time (121 answers)GPT-6.1 Sol: said 100% sure on average, right 99% of the time (2,165 answers)GPT-6.1 Sol
Jev (TypeSafe) DeepSeek V4 Flash GPT-6.1 SolA calibrated model sits on the dashed line. Each point is a confidence band (50–60%, 60–70% … 90–100%); bands with fewer than 10 answers are left out, which is why the general-purpose models have fewer points: they almost always said 90% or more. Hover a point for its count.

What stands out

How it was built

The numbers

Measure (headline items)JevDeepSeek V4 FlashGPT-6.1 Sol
Answers scored2,3772,2822,397
Accuracy86.9%94.7%95.1%
Accuracy, 95% interval84.5–89.1%93.0–96.0%93.5–96.7%
Calibration error (lower is better)0.0450.0350.021
Calibration error, 95% interval0.031–0.0680.023–0.0510.009–0.036
Brier score (lower is better)0.0960.0460.035
Gave no usable answer01150
Answered exactly 50%, so not scored for accuracy2000
Median response time (observed, not registered)0.18 s3.6 s1.6 s
Slowest 10% of answers took over (observed)0.22 s15.4 s3.8 s
Measured cost per 1,000 answers (described only)$0.03$0.15$1.18

What it does not show

Results as of Oct 10, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.