Pale strands entering from the left at scattered heights and settling into one narrow band, one of them lime, on a dark petrol-to-green gradient
Oct 6, 2026

When 90% means 90%: calibrating one-token LLM decisions

Marko Rosenmüller

Marko Rosenmüller, PhD

Technical Lead AI

TL;DR: In our previous post, we turned GLM-5.3-Flash into a decision model like TypeSafe's Jev, a hosted service that picks one of a fixed set of options and returns a probability for each, from a single token. In this post, we show how to make those probabilities reliable, so that a stated 90% means the answer is right nine times out of ten.

Our key findings:

  • Raw probabilities are overconfident. GLM-5.3-Flash states 92.7% confidence on average but is right 78.1% of the time, and it is overconfident on every dataset we tested.
  • A default temperature fixes most of it, without labels. Computed only from the number of options, it cuts the excess calibration error from 0.129 to 0.032. Out of the box, Jev is better calibrated (0.080); given the same temperature treatment, Jev reaches 0.040, putting it slightly behind. The default held up on five new tasks that nothing was tuned on.
  • With about 100 labels, calibrate() tunes the model to your task. It fits a temperature and a bias per option; the bias corrects some of the answers, so 2.0 percentage points more of them are right. It can also tell you which answers are safe to act on without review, keeping their error rate below a limit you set.

The summary compares all of them at a glance.

We also get more answers right in the first place: asking the question both before and after the state raises the share of correct answers by 1.6 percentage points.

Everything described here is now the default in privatemode-decisions and in the playground from the previous post.

Why calibration matters

A decision model returns a probability for each option, and software acts on those numbers. A typical integration acts automatically when the top option has at least 90% probability and routes everything else to a person. That rule only works if 90% means the answer is right nine times out of ten. A model that says 90% but is right only 60% of the time automates too much, and you find out only from the errors.

This property is called calibration: among all answers given with confidence p, a fraction p should be correct. Calibration is distinct from accuracy, the share of answers that are correct. An accurate model can be badly calibrated, and an inaccurate one can be well calibrated, in which case it at least tells you when not to trust it.

We measured calibration for the setup from the previous post: GLM-5.3-Flash on Privatemode, prompted with a prefilled answer and read from a single masked token. The methods are standard in machine learning. What is new is applying them to the probabilities of one token from an off-the-shelf LLM, without fine-tuning, and comparing the results with Jev on the same examples.

How we measured

We used the 29 public datasets of our benchmark, covering intent routing, sentiment, topic classification, moderation, entailment, question answering, legal text, and one set of scanned documents, with between 2 and 151 options each. We took up to 1,000 examples per dataset and split each one at random into two halves. The calibration half plays the role of the examples a user would label; all results are scored on the test half, which nothing was fitted on. Every dataset counts equally in the averages, and we ran everything twice.

The standard summary metric is the expected calibration error (ECE). It sorts the answers by stated confidence into 15 equally sized bins, compares each bin's average confidence with the share of its answers that were correct, and averages the gaps. An ECE of 0.05 means that stated confidence and the actual share of correct answers differ by 5 percentage points on average. Even a perfectly calibrated model shows some ECE on a few hundred examples, purely from sampling noise; on our test halves, that floor is about 0.05. We therefore report the excess ECE, the ECE above that floor. An excess ECE of 0 means the model is as well calibrated as the test set can show.

The calibration results use the prompt of the previous post, which puts the state before the question. The library's new default asks the question first; we describe that change at the end, where we also checked that the default temperature still fits.

Raw probabilities are overconfident

Averaged over the 28 text datasets, GLM-5.3-Flash assigns 92.7% probability to its chosen option but is right only 78.1% of the time. It is overconfident on every one of the 28 datasets, and most of all on hard tasks: on emotion classification, it is 93% confident and 60% accurate.

Stated confidence against accuracy
accuracyconfidence, rawconfidence, default temperature
patent
emotion
sst5
gnad10
amazon_reviews_de
scotus
tweet_sentiment
newsgroups20
yahoo_topics
massive_scenario_de
ledgar
toxic_conversations
banking77
tweet_offensive
massive_scenario_en
massive_intent_de
massive_intent_en
xnli_de
trec_fine
ag_news
mnli
clinc150
trec_coarse
rotten_tomatoes
rte
boolq
sst2
dbpedia_14
40%60%80%100%
Hover or tap a mark for its numbers. GLM-5.3-Flash with one token, on the test half of each of the 28 text datasets. Stated confidence is the mean probability of the chosen option; accuracy is how often that option was right, and no temperature changes it. A calibrated model puts its confidence on top of its accuracy. The ring is the probability as the model returns it; the filled dot is the same after the library’s default temperature, which depends only on the number of options and was fitted without the dataset in question (see below). Sorted by the raw gap.

The filled dots show where this post ends up without any labels. With the default temperature described below, the average gap between stated confidence and accuracy shrinks from 14.6 to 2.3 percentage points. Individual datasets still deviate: the default remains 18 points too confident on emotion and 16 on patent, two of the hardest tasks, and 8 points too cautious on mnli, which is easy for the model. Labels from your own task close that remaining gap.

The overconfidence itself is expected. LLMs are trained to predict the next token, and the fine-tuning that teaches them to follow instructions tends to make their output distributions sharper than their knowledge warrants. The probabilities we read from a single token inherit this directly.

One number fixes it: temperature scaling

The simplest remedy is temperature scaling: divide the log probabilities by a constant T and renormalize. A temperature above 1 softens the distribution, moving probability from the top option to the others. Because it never changes which option is on top, the answers, and with them the accuracy, stay exactly the same; only the stated confidence changes.

For each dataset, we fitted the temperature that best matches the labels of its calibration half. On every dataset, this brought the error on the test half down to the sampling floor. The fitted values range from 1.4 to 3.7, with a median of 2.0, and they are stable: the values fitted on our two runs differ by a factor of only 1.006.

Fitting a temperature this way, however, takes a few hundred labels per task. If a good temperature can be chosen without any labels, every decision is better calibrated out of the box, from the very first request.

A temperature without labels

The best temperature isn't arbitrary. It depends mostly on how hard a task is for the model: the model is roughly equally confident everywhere, so it is most overconfident where it is least accurate. But difficulty is unknown until you have labels, and nothing we could compute from the model's own outputs, such as its mean confidence or entropy, predicted the temperature better than a constant.

What does help is the number of options. Tasks with few options need somewhat more softening than tasks with many, and a two-parameter formula captures this: ln T = 0.962 − 0.076 · ln(options), with natural logarithms. It gives T = 2.5 for two options and T = 1.8 for 150.

Best temperature per task, by number of options
accuracy 90% and more70 to 90%below 70%
default from the option count
1234T = 1: the raw probabilities25102050150options (log scale)
Hover or tap a mark for its numbers. Each dot is one text dataset: the temperature that minimizes the calibration error on its calibration half, fitted with its labels. All of them are above 1, so every task needed softening. The line is the library’s default, log T = 0.962 − 0.076 · log(options), which needs no labels from your task. The spread around it comes mostly from how hard a task is: the darker the dot, the less accurate the model is on that dataset, and the more softening it needs.

To evaluate the formula fairly, we fitted it on all datasets except the one being scored, and repeated this for every dataset. Measured this way, without a single label from the task, it gets 71% of the way in ECE from the raw probabilities (0.149) to a temperature fitted per task (0.054); the excess ECE falls from 0.129 to 0.032. By the same measure, one temperature shared by all tasks covers 63%, and a temperature per task family (sentiment, intent, topic, and so on) 74%, provided you know which family your task belongs to.

Related datasets could inflate this result, since the benchmark includes, for example, four variants of MASSIVE. Holding those out together changes nothing (still 71%), and fitting on a random half of the tasks and testing on the other half yields 68%. The realistic worst case is an entirely new kind of task. When every dataset of the same family is held out, the formula covers 45%, which is the number to expect for a task unlike anything in the benchmark.

The library now applies this formula by default for GLM-5.3-Flash. SystemOne(..., temperature="sentiment") uses a task family's temperature instead, and temperature=1 returns the raw probabilities.

Comparison with Jev

We applied the same analysis to Jev's probabilities from the runs in the previous post, on the same examples. Out of the box, Jev is better calibrated than GLM-5.3-Flash, with an excess ECE of 0.080 compared with 0.129.

Jev rounds its probabilities to 0.01, and in 4.3% of the examples it assigns exactly 0 to the correct answer, which no temperature can change. We therefore replaced Jev's zeros with 0.005, half a rounding unit, and fitted a default temperature on Jev's own results exactly as we did for GLM-5.3-Flash. With that, Jev reaches an excess ECE of 0.040, and GLM-5.3-Flash with its default reaches 0.032, at nearly the same accuracy (78.1% compared with 77.5%). With about 500 labels and a temperature per task, the order holds: 0.006 compared with 0.019.

One caveat applies: the comparison uses the benchmark's 28 datasets, on which the method was also developed. GLM-5.3-Flash's calibration held up on five new tasks (see below), but we didn't run Jev on those.

GLM-5.3-Flash
Raw0.129
Same T for every task0.040
T from the option count0.032
T from the task family0.029
T per task~500 labels0.006
Jev
Jev, as returned0.080
Jev, own default T0.040
Jev, T per task~500 labels0.019
Hover or tap a mark for its numbers. Excess calibration error: the expected calibration error (ECE) on each dataset’s test half minus what sampling alone produces on that many examples, averaged over the 28 text datasets both systems answer. 0 means as calibrated as the test sets can show. All rows cover the same 13,005 examples, and no temperature was fitted on them. The zero-label temperatures come from the other datasets only. Jev reports probabilities rounded to 0.01, and a temperature cannot move a 0, so its zeros are set to 0.005 before one is fitted.

Averages hide where the error sits, so the reliability diagram below shows the whole picture. It groups all answers by stated confidence and plots how often each group was right. Raw GLM-5.3-Flash claims 92% confidence where it is right 60% of the time. Jev's raw probabilities are less extreme, but still 15 to 18 points too confident in the middle of the range. With the default temperature and no labels, GLM-5.3-Flash stays within about 6 points everywhere, close to what a temperature fitted on 500 labels achieves.

Share correct, by stated confidence
40%60%80%100%40%60%80%100%stated confidence
Gap to calibrated, in points
−30−20−10040%60%80%100%stated confidence
Hover or tap a mark for its numbers. All 28 text datasets’ test halves pooled, every dataset weighted equally, in 15 bins of equal size by stated confidence; each point is one bin. On the dotted line, a system is right exactly as often as it says. Default temperatures were fitted without the dataset in question and without its labels; the per-task temperature on the dataset’s calibration half. Jev’s zeros are set to 0.005 before a temperature is applied. Jev’s curves zig-zag above 85% because it rounds to 0.01: many answers share a few confidence values, and the equal-size bins split them up. Click a legend entry to hide or show its line.

Tuning to your own task

Everything so far requires no labels from you, and the defaults are the right place to start. For a decision you run at volume, though, it pays to tune the model to that specific task: have a person label a random sample of about 100 real inputs, and let calibrate(answers, labels) fit your task from them. It uses the labels for four things, each explained below: a temperature for your task, a bias per option that can correct answers, prediction sets, and a threshold for automating answers under an error bound. The figure shows how each of them improves as the number of labels grows.

What labels from your task buy
Calibration error
task temperature and bias · lower is better
0.004
at 100 labels
0.000.010.020.0302050100250500
Accuracy gained
bias per option
+2.0 pts
at 100 labels
+0+1+2+302050100250500
Options in a 90% set
prediction sets · lower is better
1.70
at 100 labels
12302050100250500
Automated at a 10% error bound
error bound
23%
at 100 labels
0%20%40%02050100250500
Hover or tap a mark for its numbers. What calibrate() returns with its defaults, by the number of labeled examples from the task (the axis is not to scale). For each dataset, that many labels were drawn at random from its calibration half, 20 times, and everything was scored on its test half. 0 labels is the default temperature alone; prediction sets and the error bound need labels. The 90% sets covered the right answer 90.9% to 92.4% of the time at every label count. At 250 labels, 27 datasets count, and at 500, the 23 with that many calibration examples.

As a rule of thumb, 20 labels already improve calibration and raise accuracy by about one percentage point. About 100 labels deliver most of the accuracy gain and make prediction sets and the error bound practical. Beyond that, the returns diminish, except for automation, which keeps improving with more labels.

The fitted calibration has three methods. apply(answer) returns the corrected answer, which is the one to use. predict_set(answer) returns the options a person should consider, and automate(answer) tells you whether to act on the answer without review.

The labels must be a random sample of your inputs. Labels collected only from cases that were escalated to a person are skewed toward hard cases and invalidate the coverage and error guarantees described below. A calibration also holds only for the model and options it was fitted on; the library rejects answers from any other.

A temperature for your task

The default temperature is fitted to the average task, not to yours. With labels, calibrate() fits one for your task, pulled toward the default with a weight equivalent to 5 examples, so that a handful of labels can't push it far. Starting at 20 labels, this beats both the default and an unconstrained fit, and with 100 or more, the pull no longer matters. On its own, the task temperature reduces the excess ECE from 0.032 with the default to 0.008 with 100 labels; combined with the bias below, to 0.004. With fewer than 10 labels, calibrate() fits only the temperature.

A bias per option

Some errors aren't about confidence at all. A model may favor one option regardless of the input, for example by predicting toxic too often. calibrate() therefore fits a bias per option alongside the temperature, again pulled toward zero so that a few labels can't push it far. For two options, this is Platt scaling. Unlike the temperature, the bias can change answers. With 100 labels, it raises accuracy by 2.0 percentage points on average, improving 18 datasets and worsening 6, none of them by more than 0.9 points. The largest gains come where the model over-predicts a class: +14.6 points on toxic_conversations and +9.5 on sst5.

Applied the same way to Jev, the bias raises Jev's accuracy by a similar amount. After calibration with 100 labels, GLM-5.3-Flash reaches 80.2% accuracy and Jev 79.6%, with excess ECEs of 0.005 and 0.014, respectively.

Prediction sets

Sometimes you don't need a single answer, only a short list that almost certainly contains it, such as the three most likely teams for a support ticket, for a person to choose from. Conformal prediction uses the labels to compute a probability cutoff such that the set of options above it contains the correct one with a chosen probability, on average over new inputs drawn like the labeled ones. This guarantee comes from the method itself, and it holds even when the model is poorly calibrated.

At a 90% target, the sets contained the correct answer 90.8% of the time, with 1.8 options on average; 64% of them contained a single option. With one cutoff for all options, though, a rare class can be covered less often than the target. The toxic class, which makes up 8% of toxic_conversations, was covered only 70% of the time. per_class=True sets a cutoff per option and raised that to 97%.

This is also where Jev's rounding becomes visible. On clinc150, with 151 options, GLM-5.3-Flash's 90% sets contain 1.07 options on average. Jev's raw probabilities give sets of all 151 options, because the correct answer is so often assigned exactly zero, and still 18 after the zeros are replaced.

Automation with an error bound

For automation, the most useful control is a bound on the error rate among the answers you act on without review. The obvious approach is to choose the probability threshold at which the observed error on your labels equals the bound. In our tests, that approach broke the bound in about 40% of the samples: on new inputs, the error among the automated answers was higher than promised. Because the threshold is tuned to the labels at hand, it inherits their luck.

calibrate(..., max_error=ε) uses Learn then Test instead. It tests candidate thresholds in order with an exact binomial test, stops at the first one that fails, and is designed to keep the error among automated answers at most ε with 90% probability. The library's version is modeled on Learn then Test rather than proven for every setting, so we validated it empirically: across all label budgets, the error on the test halves exceeded ε in at most 1.1% of the samples, well below the 10% the method allows.

The bound comes at a cost in automation. Using each dataset's full calibration half, 42% of answers could be automated at a 10% error bound, compared with 63% if the test labels had been known in hindsight. Label errors in the datasets make low bounds hard to reach at all. With only 20 labels, nothing can be certified: even 20 correct answers out of 20 don't demonstrate a 10% error rate with 90% confidence; that takes 22.

Automated under an error bound
under the boundknowing the answers
error at most 2%
6% of 31%
error at most 5%
19% of 48%
error at most 10%
42% of 63%
Hover or tap a mark for its numbers. The share of answers automated under an error bound set in the style of Learn then Test, meant to hold with 90% probability, using each dataset’s whole calibration half (about 500 labels); no dataset exceeded it on its test half. Against it, the share one could automate at that error rate if the test labels were known.

One more detail matters for anyone building this themselves. Fitting the bias and setting the cutoffs on the same labels made the answers look better than they are, and the error bound slipped on 2 of 23 datasets. calibrate() therefore sets cutoffs and thresholds on out-of-fold probabilities: each labeled answer is corrected by a fit on the other four-fifths of the labels. That preserves both the coverage and the error bound, at the cost of some automation (23% instead of 27% at a 10% bound with 100 labels). bias=False restores the temperature-only behavior if you value automation over accuracy.

A test on tasks we didn't tune on

All methods above were developed on the benchmark's 28 datasets. To test them on data they weren't developed on, we ran five tasks from domains the benchmark doesn't cover, with the measures and pass criteria fixed in advance: the topic and the sentiment of financial news posts, the field of arXiv abstracts, the study type of PubMed abstracts, and the type of GitHub issues, the last three all published in 2026. Two of the questions, topic and sentiment, also occur in the benchmark. Each task was run once, and nothing was tuned on it.

Every pre-registered criterion passed. Without calibration, the model was even more overconfident than on the known tasks, with 95.3% stated confidence at 78.8% accuracy. The default temperature reduced the excess ECE from 0.151 to 0.047 and, measured as above, got 78% of the way in ECE to a per-task temperature, against 71% on the known tasks. With 100 labels, calibrate() raised accuracy by 2.7 percentage points, with gains on every task. Its 90% prediction sets contained the correct answer 91.2% of the time with 1.5 options on average, and only 0.8% of the samples exceeded the 10% error bound.

Two results are weaker. Without labels, the arXiv and GitHub tasks were still 0.08 away from calibrated. Both are harder than the model's confidence suggests, and their own best temperatures (2.8 and 3.5) are well above the default. With 100 labels, calibrate() brings both to 0.016 or below. Automation under the error bound was also lower: 9% of the answers at a 10% bound with 100 labels, compared with 23% on the known tasks. The test stops at the first threshold that fails, and on these tasks some of the most confident answers were already wrong, several apparently because of label errors. The error bound itself held. The full report has the numbers per task.

Summary of the methods

The chart shows how much of the calibration error each method leaves, and the list what each one needs and gives you, on the benchmark's 28 text datasets:

Raw probabilitiesno labels0.129
Default temperatureno labels0.032
Task-family temperatureno labels0.029
calibrate(): temperature and bias100 labels0.004
Hover or tap a mark for its numbers. Excess calibration error of GLM-5.3-Flash, averaged over the test halves of the 28 text datasets; 0 means as calibrated as the test sets can show. The default and family temperatures need no labels from the task. The calibrate() row fits a temperature and a bias per option on 100 random labels per task, averaged over 20 draws.
  • Default temperature (no labels): computed from the number of options; better-calibrated probabilities, same answers.
  • Task-family temperature (no labels, but you name the family, e.g. temperature="sentiment"): slightly better, if your task fits the family.
  • calibrate() (about 100 labels): a temperature and a bias per option for your task; nearly calibrated probabilities, and 2.0 percentage points more answers right.
  • predict_set() (same labels): a short list of options that contains the right one 90% of the time, 1.7 options on average.
  • automate() (same labels): which answers to act on without review, with their error rate bounded; 23% of answers at a 10% bound.

In practice: start with the default, which needs nothing from you. Once a decision runs at volume, label about 100 random inputs and let calibrate() tune the model to it.

What didn't work

We also tried several other published methods, on the same examples and under the same rules. None became a default; the calibration report has the details.

  • Contextual calibration divides out the model's answer to a content-free input such as N/A. It made 24 of 28 datasets worse, because that answer is mostly real knowledge: for an empty input, neutral or other often is correct.
  • Batch calibration estimates the same bias from unlabeled traffic, assuming all options are equally common. It lowered accuracy by 0.5 percentage points on average and by 7 on toxic_conversations, where only 8% of messages are toxic.
  • Averaging over rotated option orders and PriDe target position bias, but neither improved accuracy; averaging four orders improved calibration slightly, for up to four times the requests (permutations=4).
  • Isotonic regression needs far more labels than a temperature: with 20 labels, its ECE was 0.123 against 0.066, and it only caught up at about 500.
  • The probability left outside the options doesn't flag errors: the model puts about 99% on the options whether it is right or wrong (AUROC 0.44, against 0.74 for the confidence). The library still reports it as option_mass.
  • Clustered conformal prediction had too few labels per option to group options on our datasets with many options, and made the sets larger (1.60 to 1.78 options).
  • Thermometer predicts a task's temperature from the model's internal states and would require changes to the serving stack, for a gap that 20 labels already halve.

Accuracy: ask the question first

In the previous post's prompt, the state came first, followed by the question and the options. LLMs, however, read left to right: each position can only attend to what came before it. While reading the state, the model therefore doesn't yet know what it will be asked about it. The fix is to state the question and options before the state as well:

Previous post
instructionstatequestion + optionsanswer=
Now
instructionquestion + optionsstatequestion + optionsanswer=
The previous prompt asked the question only after the state. The new one also states the question and its options before it, so the model reads the state knowing what it will be asked.

On the test halves of all 29 datasets, with two runs of each prompt, this raised mean accuracy by 1.6 percentage points, from 77.6% to 79.3%. It was better on 15 datasets, within a point on 12, and worse on 2 (Wilcoxon p = 0.002). On the 28 datasets that Jev answers, where the previous post found the two on par, GLM-5.3-Flash is now ahead: 79.8% against Jev's 77.5%, better on 16 datasets, tied on 8, and worse on 3.

The gain is concentrated where the model was unsure: +9.7 points in the least confident fifth of answers, and roughly zero elsewhere. The largest gains were on sentiment and intent tasks, such as +10.3 points on sst5 and +6.3 on toxic_conversations. It doesn't help with knowledge: on MMLU-Pro, accuracy went slightly down, from 63.1% to 61.9%. And it doesn't come from leaning more on the option names, since renaming every option to a synonym costs both prompts about the same (−7.7 and −8.1 points).

Accuracy with the question first, minus state first
-20+2+4+6+8+1015 better12 within a point2 worse
Hover or tap a mark for its numbers. GLM-5.3-Flash on the test halves of all 29 datasets, two runs of each prompt; each bar is the difference of their means. The grey band is ±1 percentage point, about what two identical runs differ by, and counts as a tie. The chance estimate is a two-sided Wilcoxon signed-rank test over the per-dataset differences (p = 0.002).

The default temperature still fits the new prompt. Refitting the formula on each prompt's answers, again leaving the scored dataset out, gives an excess ECE of 0.026 for the new prompt and 0.027 for the old one, so the new prompt is calibrated as well as the old one and needs no new constants. The shipped formula, fitted on the old prompt's answers, gives 0.016 on the new prompt and 0.023 on the old one. These numbers come from a different split of the test halves, so they don't compare directly with the 0.032 above.

The price is the question's tokens appearing twice: on average 1.6 times the input tokens, and therefore 1.6 times the cost. Up to a few dozen options, latency doesn't change measurably. With long option lists, it does: +110 ms at 77 options and +330 ms at 151.

When a request has several questions about the same state, the library asks the model once per question, and its optimize parameter chooses between two layouts. By default (optimize="accuracy"), each model call leads with its own question. With optimize="cost", every model call leads with all of the request's questions, so the calls share a prefix that the server can cache. With five questions per state, on a development set, the default raised accuracy by 2.2 percentage points and the shared layout by 0.9. The cache pays off only for long states, because Privatemode caches prefixes from about 2,300 tokens: on 2,000-token states, "cost" was the faster layout (670 ms median latency with 59% of the prompt cached, against 935 ms), while on short states it was the slower one (530 ms against 409 ms).

We also tried giving the model more to work with in the same forward pass, with filler tokens or more repetition, and nothing beat this prompt; the accuracy report has the details. Real thinking before the answer does more, but it turns milliseconds into seconds and is no longer a one-token decision.

JevBench

JevBench is an independent benchmark for Jev-like decision models, with a public leaderboard. Our system isn't on the leaderboard yet, as submissions wait in a long queue. Until then, we answered JevBench's 231 public items the same way every other system did, without tuning anything on them, and compare accuracy on those.

With one token, GLM-5.3-Flash reaches 89.4% on them (the mean of two runs), up from 88.5% with the previous post's prompt. Jev 1.13.0 reaches 86.6% on the leaderboard. Ranked by public accuracy among the decision models on the leaderboard, as of 27 September 2026, our setup would place fifth of 90, behind one Gemma-based service at 92.8% and three systems at 89.6%, one item ahead. The leaderboard also lists general-purpose LLMs as baselines; those that reason before answering reach up to 99.6%, at seconds per decision, and are left out of the figure.

JevBench decision models, public items · as of 27 September 2026
GLM-5.3-Flash, one tokenJev 1.13.0other decision models
1Autoloops – Gemma 4 31B IT92.8%
2Plumb-4B (crh225, JevK5 v0.2 + LoRA)89.6%
3NInfer Qwen3.8-Flash-Next mixed89.6%
4JevOne89.6%
5GLM-5.3-Flash on Privatemode, one token89.4%
6Jev-Omni (akhilaaa3, Gemma-4-12B merged)88.7%
7swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4)88.7%
8OpenJev (thinking, BF16)88.7%
9Cygnet (blockbrain, frozen Gemma-4-12B-it)87.9%
10djev (thinking)87.4%
11reflex-27b (Qwen3.8-27B)87.0%
12Jev 1.13.0 (TypeSafe AI)86.6%
13SimpleJev Qwen3.8-27B86.6%
14Instinct (ZooWork, Qwen3.8-27B)86.6%
15Imajev-4B86.1%
16LitJev (Qwen3.8-27B)86.1%
17Winnow-12B Q885.7%
18JevK5 v0.2.085.3%
19openjev-sglang (Qwen3.6-35B-A3B on SGLang)85.3%
20classifier.dev (fast tier)85.3%
and 70 more80%90%100%
Hover or tap a mark for its numbers. Accuracy on JevBench’s 231 public items, as of 27 September 2026 (JevBench v1.4.2.2). The other systems are the board’s 89 decision models, the top 19 shown; the general-purpose LLMs it lists as baselines are left out. Ours was measured by us with the library’s current prompt, the mean of two runs, and is not part of the official board. The axis starts at 80%. The official JevBench score also weighs held-out items, calibration, speed, and cost, and can’t be computed from the public items alone.

This comparison comes with caveats. Public items can be trained on or selected against, which is why JevBench also scores held-out items. On those, every system on the leaderboard is far less accurate (Jev: 36.7%), so public accuracy doesn't predict the official rank. JevBench's items are also closer to rules and judgment calls than our classification datasets. On MMLU-Pro, a knowledge and reasoning test, the picture reverses: Jev reaches a published 82.9%, and GLM-5.3-Flash in a single pass 61.9%. For decisions that require knowledge or several reasoning steps, a single token from a Flash-class model is not enough.

Try it

The playground in the previous post now uses everything described here: the question-first prompt and the default temperature for GLM-5.3-Flash. Each answer shows the temperature its probabilities were scaled by. The trick question there remains a good reminder of what calibration can't do: the model answers it wrong in one token, and calibration only makes it less sure of the wrong answer.

One more change since the previous post is easy to miss: the prefill is now answer= instead of choice_index:. After a colon, the model would rather write a space than a digit, and the chat template strips trailing spaces, so less than 0.1% of the probability landed on the options. With answer=, nearly all of it does. Accuracy stayed the same, and all numbers in this post use answer=.

Limitations

  • Five held-out tasks are few. The held-out test confirms the method on tasks it wasn't tuned on, but five tasks from four new domains are a small sample, and Jev wasn't run on them.
  • One image dataset. Scanned documents (rvl_cdip) are overconfident too, and the text formula happens to give nearly their own best temperature (2.12 compared with 2.14). A single dataset is thin evidence.
  • Label errors. Some of the remaining "errors" are wrong labels. We had Claude review all 80 disputed banking77 examples: 31% were clear label errors, and 50% had two defensible answers. These judgments haven't been checked by hand yet.
  • Other models need their own numbers. Kimi K2.6 follows the same recipe, and its default temperatures ship with the library. GLM-5.3 (not Flash) doesn't: its best temperature doesn't depend on the number of options, and after answer=, it puts about a third of its probability elsewhere, mostly on a space.

Build it yourself

Both the library and the benchmark are open source. The benchmark repository contains the full calibration report in three parts, with every table in this post and many more, the code that produced them, the raw calibration runs, and tests that check the library's calibration code against the report. The library's README explains which guarantee holds when, and its AGENTS.md tells coding agents what an implementation has to get right.

edgelesssys/privatemode-decisionsThe Python library: token oracle, prompt, masking and renormalization, against any vLLM-backed endpoint. edgelesssys/privatemode-decisions-benchmarkThe benchmark: methodology, frozen dataset specs, harness and aggregation. Every number in this post can be recomputed from it.

With Privatemode, these decisions are protected by confidential computing: your data stays encrypted in memory even during processing, and the client verifies the deployment's attestation report before sending anything. You can read more on our security page.

Build it on Privatemode

Create an account, then give this post to Claude Code, Codex, or the coding tool of your choice. Together with the GitHub repository, it has everything needed to build typed decisions into your own code, protected by confidential computing.

Create an account

Articles

Further reading

Explore other articles

Three nested outlines around a lit core, one dotted line running out to a small node, on a petrol-to-mint gradient

Oct 2, 2026

OpenAI Private Intelligence: what it promises, and what you can verify today

OpenAI now says AI inference belongs in confidential computing. What the DevDay announcement contains, what it leaves open, and three questions that separate a verifiable property from a promise.

A plume of pale filaments rising from a single point, one of them mint and reaching furthest, on a teal-to-steel gradient

Sep 24, 2026

Turn GLM-5.3-Flash into a Jev-like System One model

Typed decisions in a single forward pass, about as accurate as Jev and faster from Europe. With a playground and a benchmark on 29 datasets.