Jev now has competitors. Does any decision model know when it is wrong?

Jev, OpenAI, AWS, Perplexity: Do Decision Models Know When They're Wrong?

Oct 11, 2026

A month ago, decision models were a single product: TypeSafe's Jev. Now there are at least seven, from OpenAI, Perplexity, the Strands Agents project that came out of AWS, Fastino, Convai Innovations and an independent developer. Each one answers a yes/no question with a probability, and the pitch is that the probability can be trusted. Every published claim for them so far is the vendor's own. We gave all seven the same 799 yes/no decisions with known answers, items we have never published, so no model can have been trained on them, under a plan registered before the run. One model, Perplexity's open pplx-decider, was both the most accurate (94.7%) and the only one whose confidence matched its accuracy under our registered rule. OpenAI's Decisions API and Jev were too close to call, and statistically tied with each other. Three models did not support a calibration claim, and two of them, Strands Decider and Laya, scored below 67.7%, the score of always giving the most common answer.

Full paper: SILT-RP-010 · DOI 10.5281/zenodo.23293732 — the complete method, every table and the registration.

Accuracy and calibration on the same 799 questionsPerplexityOpenAI DecisionsJev (TypeSafe)OpenDeciderFastinoStrands DeciderLayaAccuracyPerplexity · Accuracy: 94.7%94.7%OpenAI Decisions · Accuracy: 88.0%88.0%Jev (TypeSafe) · Accuracy: 86.4%86.4%OpenDecider · Accuracy: 84.0%84.0%Fastino · Accuracy: 69.7%69.7%Strands Decider · Accuracy: 66.3%66.3%Laya · Accuracy: 64.8%64.8%Calibration error (lower is better)Perplexity · Calibration error (lower is better): 0.0260.026OpenAI Decisions · Calibration error (lower is better): 0.0340.034Jev (TypeSafe) · Calibration error (lower is better): 0.0520.052OpenDecider · Calibration error (lower is better): 0.0790.079Fastino · Calibration error (lower is better): 0.0820.082Strands Decider · Calibration error (lower is better): 0.1290.129Laya · Calibration error (lower is better): 0.0510.051
Each panel has its own scale. The three models in colour are the ones drawn in the next chart. Accuracy is on the headline items; always giving the most common answer would score 67.7%. Calibration error is the registered measure: 0.05 or under, with the whole interval under it, passes. The numbers and intervals are in the table below.
Stated confidence against share correct, headline items50%60%70%80%90%100%30%40%50%60%70%80%90%100%How sure the model said it wasHow often it was rightdashed line = perfectly calibratedbelow it = overconfidentPerplexity: said 55% sure on average, right 62% of the time (39 answers)Perplexity: said 65% sure on average, right 89% of the time (27 answers)Perplexity: said 76% sure on average, right 72% of the time (87 answers)Perplexity: said 86% sure on average, right 65% of the time (129 answers)Perplexity: said 99% sure on average, right 98% of the time (2,109 answers)PerplexityJev (TypeSafe): said 56% sure on average, right 51% of the time (249 answers)Jev (TypeSafe): said 65% sure on average, right 71% of the time (356 answers)Jev (TypeSafe): said 75% sure on average, right 85% of the time (349 answers)Jev (TypeSafe): said 85% sure on average, right 91% of the time (399 answers)Jev (TypeSafe): said 96% sure on average, right 99% of the time (1,036 answers)Jev (TypeSafe)Laya: said 54% sure on average, right 55% of the time (504 answers)Laya: said 66% sure on average, right 58% of the time (720 answers)Laya: said 75% sure on average, right 72% of the time (717 answers)Laya: said 84% sure on average, right 78% of the time (417 answers)Laya: said 93% sure on average, right 50% of the time (36 answers)Laya
Perplexity Jev (TypeSafe) LayaA calibrated model sits on the dashed line. Each point is a confidence band (50–60%, 60–70% … 90–100%); bands with fewer than 10 answers are left out. Three models are shown, the most accurate, Jev and the least accurate; every model's bands are in the full results. Hover a point for its count.

What stands out

How it was built

The numbers

Model (headline items)AccuracyAccuracy, 95% intervalCalibration error (lower is better)Calibration error, 95% intervalBrier score (lower is better)Registered verdict
Perplexity pplx-decider (open, 27B)94.7%93.1–96.4%0.0260.014–0.0440.044Consistent with the calibration claim
OpenAI Decisions API (gpt-6-luna)88.0%85.6–90.1%0.0340.025–0.0620.098Inconclusive
Jev (TypeSafe)86.4%84.0–88.7%0.0520.034–0.0720.096Inconclusive
OpenDecider (independent, open)84.0%81.3–86.6%0.0790.062–0.1050.123Not supported on these items
Fastino GLiNER2.5-Decide (open)69.7%66.8–72.7%0.0820.059–0.1150.191Not supported on these items
Strands Decider (Strands Agents, open)66.3%62.8–69.8%0.1290.100–0.1580.188Not supported on these items
Laya (Convai Innovations, open)64.8%61.5–68.1%0.0510.029–0.0880.224Inconclusive
For context: always the most common answer67.7%———0.219—
For context: DeepSeek V4 Flash (general model)94.7%93.0–96.0%0.0350.023–0.0510.046No verdict (context)
For context: GPT-6.1 Sol (general model)95.1%93.5–96.7%0.0210.009–0.0360.035No verdict (context)

Also registered: how often 'yes' was right, by stated probability

Our registration promised this finer table alongside the five confidence bands. For each tenth of stated probability of yes, it shows how often the answer really was yes, with the number of answers in brackets. A calibrated model's share sits inside its row's range (for example, 70–80% in the 0.7–0.8 row). Rows with fewer than 10 answers show a dash. Laya never said below 0.4, and Fastino's answers under 0.1 were yes 73% of the time.

Model said P(yes)PerplexityOpenAIJevOpenDeciderFastinoStrandsLaya
0.0–0.14% (687)15% (516)1% (488)0% (27)73% (78)0% (30)—
0.1–0.240% (90)20% (153)25% (142)13% (180)32% (114)31% (78)—
0.2–0.344% (48)22% (81)43% (90)18% (183)24% (438)86% (66)—
0.3–0.4—61% (54)54% (173)21% (102)42% (285)69% (156)—
0.4–0.5100% (12)52% (75)86% (146)42% (201)77% (429)43% (105)76% (135)
0.5–0.691% (33)79% (117)93% (123)51% (231)81% (273)29% (330)67% (372)
0.6–0.7100% (18)83% (54)96% (182)75% (216)88% (96)53% (570)58% (720)
0.7–0.892% (39)86% (84)97% (252)87% (432)96% (243)88% (552)72% (717)
0.8–0.977% (39)80% (105)99% (271)96% (471)98% (360)100% (492)78% (417)
0.9–1.099% (1,422)98% (1,155)99% (530)100% (354)100% (81)100% (18)50% (36)

Also registered: ambiguous items, cost and cut-off inputs

Sixty items were written so that the answer cannot be known from the description. They are never scored; what matters is whether a model marks them as uncertain. Cost is what each hosted API reported for the whole run; the open models ran on our own hardware or a rented GPU. Inputs longer than a model's window were cut by that model, as shipped. Jev, OpenAI and Fastino do not report whether they cut an input, so we could not measure it for them.

MeasurePerplexityOpenAIJevOpenDeciderFastinoStrandsLaya
Ambiguous items answered between 0.3 and 0.728%65%78%60%88%43%47%
Ambiguous items: mean P(yes)0.570.510.400.650.380.690.71
Hosted cost, whole run (3,297 answers)—$0.159$0.086————
Answers whose input was cut0%not reportednot reported0%not reported0%3.0%

What it does not show

Results as of Oct 11, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.