Jev now has competitors. Does any decision model know when it is wrong?
Jev, OpenAI, AWS, Perplexity: Do Decision Models Know When They're Wrong?
Oct 11, 2026
A month ago, decision models were a single product: TypeSafe's Jev. Now there are at least seven, from OpenAI, Perplexity, the Strands Agents project that came out of AWS, Fastino, Convai Innovations and an independent developer. Each one answers a yes/no question with a probability, and the pitch is that the probability can be trusted. Every published claim for them so far is the vendor's own. We gave all seven the same 799 yes/no decisions with known answers, items we have never published, so no model can have been trained on them, under a plan registered before the run. One model, Perplexity's open pplx-decider, was both the most accurate (94.7%) and the only one whose confidence matched its accuracy under our registered rule. OpenAI's Decisions API and Jev were too close to call, and statistically tied with each other. Three models did not support a calibration claim, and two of them, Strands Decider and Laya, scored below 67.7%, the score of always giving the most common answer.
Each panel has its own scale. The three models in colour are the ones drawn in the next chart. Accuracy is on the headline items; always giving the most common answer would score 67.7%. Calibration error is the registered measure: 0.05 or under, with the whole interval under it, passes. The numbers and intervals are in the table below. Perplexity Jev (TypeSafe) LayaA calibrated model sits on the dashed line. Each point is a confidence band (50–60%, 60–70% … 90–100%); bands with fewer than 10 answers are left out. Three models are shown, the most accurate, Jev and the least accurate; every model's bands are in the full results. Hover a point for its count.
What stands out
Calibration, the registered question: one model of seven passed. Perplexity's pplx-decider had a calibration error of 0.026, with an interval (0.014 to 0.044) entirely under our line of 0.05. Three models had intervals entirely above the line (OpenDecider, Fastino and Strands Decider), so their calibration claims are not supported on these items. Jev, OpenAI's Decisions API and Laya were inconclusive.
Accuracy against Jev, the second registered question, on the same items and answers: Perplexity's model was 8.1 points more accurate (interval 6.1 to 10.1). OpenAI's Decisions API was statistically tied with Jev (+0.6, interval −2.0 to +3.1), and so, narrowly, was OpenDecider (−2.8, interval −5.8 to +0.2). Fastino (−16.5), Strands Decider (−20.1) and Laya (−21.6) were far behind.
Below a baseline that needs no model: always giving the most common answer scores 67.7% on these items. Strands Decider scored 66.3% and Laya 64.8%. Fastino scored 69.7%, just above it.
Harmless commands read as destructive: asked whether read-only commands such as git status would destroy data, Laya answered wrongly on all 300 of those answers and Strands Decider on 94% of them. The other five models were right on 97% to 100% of them.
Confidently wrong at the top: when Laya reported 90% confidence or more, it was right half the time (36 answers). When Fastino did, it was right 64% of the time (159 answers). Perplexity's model was right 98% of the time at that confidence, on 2,109 answers.
Contradicting themselves: asked a question and its exact negation, a coherent model's two probabilities of yes add up to 1. Strands Decider's averaged 1.31 and Laya's 1.44, leaning toward yes to both; OpenAI's Decisions API (0.74) and Fastino (0.72) leaned toward no to both. Perplexity's averaged 0.98.
Arithmetic is a coin flip for nearly every decision model: six of the seven, Jev included, scored exactly 50% on the arithmetic items, and OpenAI's Decisions API 63%, while both general-purpose models scored above 95%. This set is reported separately and is not in the headline.
Same answer every time: all five open models, and OpenAI's Decisions API, gave exactly the same probability on all three askings of every item. Jev did so on 27% of items. Its accuracy this time (86.4%) was close to our first Jev study a day earlier (86.9%, same version, jev-1.13.0); both studies found its calibration inconclusive.
Speed, observed, not registered: the four models on our own GPU answered in a median of 8 to 29 milliseconds. Jev's API took 0.15 seconds and OpenAI's 0.17, including the network. Perplexity's model took 0.21 seconds on the rented A100.
How it was built
The items are the ones from our Jev study, unchanged and still unpublished: every question is a yes/no decision whose answer was fixed before any model saw it, by a check of a real workspace, a deterministic simulator or arithmetic. No answer key comes from another model's judgment.
Two kinds of item count toward the headline (799 items): questions about real AI coding-agent runs from our Harness Battery, and questions asking whether a shell command, in a described workspace, would permanently destroy data.
A third set targets weak spots that vendors document (arithmetic, counting, negation, distracting context and others). It is reported separately and never counted in the headline. So are 60 items written to be genuinely ambiguous.
Every model was tested as shipped: its maker's own interface, its default checkpoint and its own saved calibration. We did no tuning on our data. Laya's model card tells users to recalibrate it first because it 'ships over-confident'; we did not do that for it, because a buyer who downloads it gets it as shipped.
Jev and OpenAI's Decisions API were called over their APIs. Four open models ran on our own RTX 5090. Perplexity's needs about 50 GB of GPU memory, more than the 5090 has, so it ran on a rented 80 GB A100. That is itself a finding for anyone planning to run it.
Six other candidates were found and left out by a gate fixed before the run, with the reason recorded: kouhxp/gutsy (under the gate's 1,000 downloads); the openjev models (community conversions of a commercial model, under a non-commercial licence); JevK5, NanoJev and jevlike (derivatives, a template for building one, and a training kit with no weights); and Fastino's hosted GLiDE API (it needed a separate account, and Fastino is represented by its official open model).
OpenAI's Decisions API has no published contract. We read its request format from the endpoint's own error messages and one test call, then fixed one adaptation for every item before the run: the answer criteria go into its instructions, because it has no separate field for them.
Every item was answered three times by each model: 23,079 answers in all. The models ran one at a time.
The test, the pass line (an expected calibration error of at most 0.05), the wording of each possible verdict and the comparison with Jev were registered on the Open Science Framework before the run (osf.io/xh7qu), and the analysis was run once. Two general-purpose models from our Jev study, DeepSeek V4 Flash and GPT-6.1 Sol, answered the same items there and are shown for context only, with no verdict.
The numbers
Model (headline items)
Accuracy
Accuracy, 95% interval
Calibration error (lower is better)
Calibration error, 95% interval
Brier score (lower is better)
Registered verdict
Perplexity pplx-decider (open, 27B)
94.7%
93.1–96.4%
0.026
0.014–0.044
0.044
Consistent with the calibration claim
OpenAI Decisions API (gpt-6-luna)
88.0%
85.6–90.1%
0.034
0.025–0.062
0.098
Inconclusive
Jev (TypeSafe)
86.4%
84.0–88.7%
0.052
0.034–0.072
0.096
Inconclusive
OpenDecider (independent, open)
84.0%
81.3–86.6%
0.079
0.062–0.105
0.123
Not supported on these items
Fastino GLiNER2.5-Decide (open)
69.7%
66.8–72.7%
0.082
0.059–0.115
0.191
Not supported on these items
Strands Decider (Strands Agents, open)
66.3%
62.8–69.8%
0.129
0.100–0.158
0.188
Not supported on these items
Laya (Convai Innovations, open)
64.8%
61.5–68.1%
0.051
0.029–0.088
0.224
Inconclusive
For context: always the most common answer
67.7%
—
—
—
0.219
—
For context: DeepSeek V4 Flash (general model)
94.7%
93.0–96.0%
0.035
0.023–0.051
0.046
No verdict (context)
For context: GPT-6.1 Sol (general model)
95.1%
93.5–96.7%
0.021
0.009–0.036
0.035
No verdict (context)
Also registered: how often 'yes' was right, by stated probability
Our registration promised this finer table alongside the five confidence bands. For each tenth of stated probability of yes, it shows how often the answer really was yes, with the number of answers in brackets. A calibrated model's share sits inside its row's range (for example, 70–80% in the 0.7–0.8 row). Rows with fewer than 10 answers show a dash. Laya never said below 0.4, and Fastino's answers under 0.1 were yes 73% of the time.
Model said P(yes)
Perplexity
OpenAI
Jev
OpenDecider
Fastino
Strands
Laya
0.0–0.1
4% (687)
15% (516)
1% (488)
0% (27)
73% (78)
0% (30)
—
0.1–0.2
40% (90)
20% (153)
25% (142)
13% (180)
32% (114)
31% (78)
—
0.2–0.3
44% (48)
22% (81)
43% (90)
18% (183)
24% (438)
86% (66)
—
0.3–0.4
—
61% (54)
54% (173)
21% (102)
42% (285)
69% (156)
—
0.4–0.5
100% (12)
52% (75)
86% (146)
42% (201)
77% (429)
43% (105)
76% (135)
0.5–0.6
91% (33)
79% (117)
93% (123)
51% (231)
81% (273)
29% (330)
67% (372)
0.6–0.7
100% (18)
83% (54)
96% (182)
75% (216)
88% (96)
53% (570)
58% (720)
0.7–0.8
92% (39)
86% (84)
97% (252)
87% (432)
96% (243)
88% (552)
72% (717)
0.8–0.9
77% (39)
80% (105)
99% (271)
96% (471)
98% (360)
100% (492)
78% (417)
0.9–1.0
99% (1,422)
98% (1,155)
99% (530)
100% (354)
100% (81)
100% (18)
50% (36)
Also registered: ambiguous items, cost and cut-off inputs
Sixty items were written so that the answer cannot be known from the description. They are never scored; what matters is whether a model marks them as uncertain. Cost is what each hosted API reported for the whole run; the open models ran on our own hardware or a rented GPU. Inputs longer than a model's window were cut by that model, as shipped. Jev, OpenAI and Fastino do not report whether they cut an input, so we could not measure it for them.
Measure
Perplexity
OpenAI
Jev
OpenDecider
Fastino
Strands
Laya
Ambiguous items answered between 0.3 and 0.7
28%
65%
78%
60%
88%
43%
47%
Ambiguous items: mean P(yes)
0.57
0.51
0.40
0.65
0.38
0.69
0.71
Hosted cost, whole run (3,297 answers)
—
$0.159
$0.086
—
—
—
—
Answers whose input was cut
0%
not reported
not reported
0%
not reported
0%
3.0%
What it does not show
Seven separate verdicts are reported without correction. At 95% each, the chance that at least one lands on the wrong side by chance alone is about 30%.
The 0.05 line is ours. No vendor publishes a calibration target, so a different reasonable line could give a different verdict.
Every model was tested as shipped. Laya's card tells users to recalibrate before use; a recalibrated Laya could do better, and that would be a different test.
Inputs longer than a model's window were cut by that model, as shipped: 3% of Laya's answers were affected. We measured this and did not fix it.
Fastino's model returns a label with a confidence. We converted that to a probability of yes with one fixed rule, set before the run.
OpenAI's Decisions API refused to answer 3 of 3,297 askings (all about the same harmless git status item). Under the registered rule a refusal is no decision, not a wrong answer. When its response format surprised our software the first time, the run stopped instead of guessing; we logged that as a deviation.
These are our items, about coding-agent work and destructive commands. They are not any vendor's own workflows, and we did not test price or speed claims. Response times are observations: the local models ran on our hardware and the hosted ones over the internet, so the two groups are never compared.
We applied our own entry gate unevenly in one case. The gate admitted a model that was either its maker's official release or past 1,000 downloads. OpenDecider, an independent developer's own release with fewer downloads, was admitted as its maker's release; gutsy, also an independent developer's own release, was excluded on downloads. Read the same way, gutsy would have qualified. Our registration forbids revisiting a gate once data exist, so we report the inconsistency instead of changing it.
The detailed findings beyond the two registered questions (confidence bands, the weak-spot set, determinism and speed) are reported as observations. All deviations from the registered plan are filed with the registration.
Run 10–11 October 2026. We report what the models did; we say nothing about intent.
Results as of Oct 11, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: