When 90% means 90%: calibrating one-token LLM decisions

Marko Rosenmüller, PhD
Technical Lead AI


Marko Rosenmüller, PhD
Technical Lead AI
TL;DR: In our previous post, we turned GLM-5.3-Flash into a decision model like TypeSafe's Jev, a hosted service that picks one of a fixed set of options and returns a probability for each, from a single token. In this post, we show how to make those probabilities reliable, so that a stated 90% means the answer is right nine times out of ten.
Our key findings:
calibrate() tunes the model to your task. It fits a temperature and a bias per option; the bias corrects some of the answers, so 2.0 percentage points more of them are right. It can also tell you which answers are safe to act on without review, keeping their error rate below a limit you set.The summary compares all of them at a glance.
We also get more answers right in the first place: asking the question both before and after the state raises the share of correct answers by 1.6 percentage points.
Everything described here is now the default in privatemode-decisions and in the playground from the previous post.
A decision model returns a probability for each option, and software acts on those numbers. A typical integration acts automatically when the top option has at least 90% probability and routes everything else to a person. That rule only works if 90% means the answer is right nine times out of ten. A model that says 90% but is right only 60% of the time automates too much, and you find out only from the errors.
This property is called calibration: among all answers given with confidence p, a fraction p should be correct. Calibration is distinct from accuracy, the share of answers that are correct. An accurate model can be badly calibrated, and an inaccurate one can be well calibrated, in which case it at least tells you when not to trust it.
We measured calibration for the setup from the previous post: GLM-5.3-Flash on Privatemode, prompted with a prefilled answer and read from a single masked token. The methods are standard in machine learning. What is new is applying them to the probabilities of one token from an off-the-shelf LLM, without fine-tuning, and comparing the results with Jev on the same examples.
We used the 29 public datasets of our benchmark, covering intent routing, sentiment, topic classification, moderation, entailment, question answering, legal text, and one set of scanned documents, with between 2 and 151 options each. We took up to 1,000 examples per dataset and split each one at random into two halves. The calibration half plays the role of the examples a user would label; all results are scored on the test half, which nothing was fitted on. Every dataset counts equally in the averages, and we ran everything twice.
The standard summary metric is the expected calibration error (ECE). It sorts the answers by stated confidence into 15 equally sized bins, compares each bin's average confidence with the share of its answers that were correct, and averages the gaps. An ECE of 0.05 means that stated confidence and the actual share of correct answers differ by 5 percentage points on average. Even a perfectly calibrated model shows some ECE on a few hundred examples, purely from sampling noise; on our test halves, that floor is about 0.05. We therefore report the excess ECE, the ECE above that floor. An excess ECE of 0 means the model is as well calibrated as the test set can show.
The calibration results use the prompt of the previous post, which puts the state before the question. The library's new default asks the question first; we describe that change at the end, where we also checked that the default temperature still fits.
Averaged over the 28 text datasets, GLM-5.3-Flash assigns 92.7% probability to its chosen option but is right only 78.1% of the time. It is overconfident on every one of the 28 datasets, and most of all on hard tasks: on emotion classification, it is 93% confident and 60% accurate.
The filled dots show where this post ends up without any labels. With the default temperature described below, the average gap between stated confidence and accuracy shrinks from 14.6 to 2.3 percentage points. Individual datasets still deviate: the default remains 18 points too confident on emotion and 16 on patent, two of the hardest tasks, and 8 points too cautious on mnli, which is easy for the model. Labels from your own task close that remaining gap.
The overconfidence itself is expected. LLMs are trained to predict the next token, and the fine-tuning that teaches them to follow instructions tends to make their output distributions sharper than their knowledge warrants. The probabilities we read from a single token inherit this directly.
The simplest remedy is temperature scaling: divide the log probabilities by a constant T and renormalize. A temperature above 1 softens the distribution, moving probability from the top option to the others. Because it never changes which option is on top, the answers, and with them the accuracy, stay exactly the same; only the stated confidence changes.
For each dataset, we fitted the temperature that best matches the labels of its calibration half. On every dataset, this brought the error on the test half down to the sampling floor. The fitted values range from 1.4 to 3.7, with a median of 2.0, and they are stable: the values fitted on our two runs differ by a factor of only 1.006.
Fitting a temperature this way, however, takes a few hundred labels per task. If a good temperature can be chosen without any labels, every decision is better calibrated out of the box, from the very first request.
The best temperature isn't arbitrary. It depends mostly on how hard a task is for the model: the model is roughly equally confident everywhere, so it is most overconfident where it is least accurate. But difficulty is unknown until you have labels, and nothing we could compute from the model's own outputs, such as its mean confidence or entropy, predicted the temperature better than a constant.
What does help is the number of options. Tasks with few options need somewhat more softening than tasks with many, and a two-parameter formula captures this: ln T = 0.962 − 0.076 · ln(options), with natural logarithms. It gives T = 2.5 for two options and T = 1.8 for 150.
To evaluate the formula fairly, we fitted it on all datasets except the one being scored, and repeated this for every dataset. Measured this way, without a single label from the task, it gets 71% of the way in ECE from the raw probabilities (0.149) to a temperature fitted per task (0.054); the excess ECE falls from 0.129 to 0.032. By the same measure, one temperature shared by all tasks covers 63%, and a temperature per task family (sentiment, intent, topic, and so on) 74%, provided you know which family your task belongs to.
Related datasets could inflate this result, since the benchmark includes, for example, four variants of MASSIVE. Holding those out together changes nothing (still 71%), and fitting on a random half of the tasks and testing on the other half yields 68%. The realistic worst case is an entirely new kind of task. When every dataset of the same family is held out, the formula covers 45%, which is the number to expect for a task unlike anything in the benchmark.
The library now applies this formula by default for GLM-5.3-Flash. SystemOne(..., temperature="sentiment") uses a task family's temperature instead, and temperature=1 returns the raw probabilities.
We applied the same analysis to Jev's probabilities from the runs in the previous post, on the same examples. Out of the box, Jev is better calibrated than GLM-5.3-Flash, with an excess ECE of 0.080 compared with 0.129.
Jev rounds its probabilities to 0.01, and in 4.3% of the examples it assigns exactly 0 to the correct answer, which no temperature can change. We therefore replaced Jev's zeros with 0.005, half a rounding unit, and fitted a default temperature on Jev's own results exactly as we did for GLM-5.3-Flash. With that, Jev reaches an excess ECE of 0.040, and GLM-5.3-Flash with its default reaches 0.032, at nearly the same accuracy (78.1% compared with 77.5%). With about 500 labels and a temperature per task, the order holds: 0.006 compared with 0.019.
One caveat applies: the comparison uses the benchmark's 28 datasets, on which the method was also developed. GLM-5.3-Flash's calibration held up on five new tasks (see below), but we didn't run Jev on those.
Averages hide where the error sits, so the reliability diagram below shows the whole picture. It groups all answers by stated confidence and plots how often each group was right. Raw GLM-5.3-Flash claims 92% confidence where it is right 60% of the time. Jev's raw probabilities are less extreme, but still 15 to 18 points too confident in the middle of the range. With the default temperature and no labels, GLM-5.3-Flash stays within about 6 points everywhere, close to what a temperature fitted on 500 labels achieves.
Everything so far requires no labels from you, and the defaults are the right place to start. For a decision you run at volume, though, it pays to tune the model to that specific task: have a person label a random sample of about 100 real inputs, and let calibrate(answers, labels) fit your task from them. It uses the labels for four things, each explained below: a temperature for your task, a bias per option that can correct answers, prediction sets, and a threshold for automating answers under an error bound. The figure shows how each of them improves as the number of labels grows.
As a rule of thumb, 20 labels already improve calibration and raise accuracy by about one percentage point. About 100 labels deliver most of the accuracy gain and make prediction sets and the error bound practical. Beyond that, the returns diminish, except for automation, which keeps improving with more labels.
The fitted calibration has three methods. apply(answer) returns the corrected answer, which is the one to use. predict_set(answer) returns the options a person should consider, and automate(answer) tells you whether to act on the answer without review.
The labels must be a random sample of your inputs. Labels collected only from cases that were escalated to a person are skewed toward hard cases and invalidate the coverage and error guarantees described below. A calibration also holds only for the model and options it was fitted on; the library rejects answers from any other.
The default temperature is fitted to the average task, not to yours. With labels, calibrate() fits one for your task, pulled toward the default with a weight equivalent to 5 examples, so that a handful of labels can't push it far. Starting at 20 labels, this beats both the default and an unconstrained fit, and with 100 or more, the pull no longer matters. On its own, the task temperature reduces the excess ECE from 0.032 with the default to 0.008 with 100 labels; combined with the bias below, to 0.004. With fewer than 10 labels, calibrate() fits only the temperature.
Some errors aren't about confidence at all. A model may favor one option regardless of the input, for example by predicting toxic too often. calibrate() therefore fits a bias per option alongside the temperature, again pulled toward zero so that a few labels can't push it far. For two options, this is Platt scaling. Unlike the temperature, the bias can change answers. With 100 labels, it raises accuracy by 2.0 percentage points on average, improving 18 datasets and worsening 6, none of them by more than 0.9 points. The largest gains come where the model over-predicts a class: +14.6 points on toxic_conversations and +9.5 on sst5.
Applied the same way to Jev, the bias raises Jev's accuracy by a similar amount. After calibration with 100 labels, GLM-5.3-Flash reaches 80.2% accuracy and Jev 79.6%, with excess ECEs of 0.005 and 0.014, respectively.
Sometimes you don't need a single answer, only a short list that almost certainly contains it, such as the three most likely teams for a support ticket, for a person to choose from. Conformal prediction uses the labels to compute a probability cutoff such that the set of options above it contains the correct one with a chosen probability, on average over new inputs drawn like the labeled ones. This guarantee comes from the method itself, and it holds even when the model is poorly calibrated.
At a 90% target, the sets contained the correct answer 90.8% of the time, with 1.8 options on average; 64% of them contained a single option. With one cutoff for all options, though, a rare class can be covered less often than the target. The toxic class, which makes up 8% of toxic_conversations, was covered only 70% of the time. per_class=True sets a cutoff per option and raised that to 97%.
This is also where Jev's rounding becomes visible. On clinc150, with 151 options, GLM-5.3-Flash's 90% sets contain 1.07 options on average. Jev's raw probabilities give sets of all 151 options, because the correct answer is so often assigned exactly zero, and still 18 after the zeros are replaced.
For automation, the most useful control is a bound on the error rate among the answers you act on without review. The obvious approach is to choose the probability threshold at which the observed error on your labels equals the bound. In our tests, that approach broke the bound in about 40% of the samples: on new inputs, the error among the automated answers was higher than promised. Because the threshold is tuned to the labels at hand, it inherits their luck.
calibrate(..., max_error=ε) uses Learn then Test instead. It tests candidate thresholds in order with an exact binomial test, stops at the first one that fails, and is designed to keep the error among automated answers at most ε with 90% probability. The library's version is modeled on Learn then Test rather than proven for every setting, so we validated it empirically: across all label budgets, the error on the test halves exceeded ε in at most 1.1% of the samples, well below the 10% the method allows.
The bound comes at a cost in automation. Using each dataset's full calibration half, 42% of answers could be automated at a 10% error bound, compared with 63% if the test labels had been known in hindsight. Label errors in the datasets make low bounds hard to reach at all. With only 20 labels, nothing can be certified: even 20 correct answers out of 20 don't demonstrate a 10% error rate with 90% confidence; that takes 22.
One more detail matters for anyone building this themselves. Fitting the bias and setting the cutoffs on the same labels made the answers look better than they are, and the error bound slipped on 2 of 23 datasets. calibrate() therefore sets cutoffs and thresholds on out-of-fold probabilities: each labeled answer is corrected by a fit on the other four-fifths of the labels. That preserves both the coverage and the error bound, at the cost of some automation (23% instead of 27% at a 10% bound with 100 labels). bias=False restores the temperature-only behavior if you value automation over accuracy.
All methods above were developed on the benchmark's 28 datasets. To test them on data they weren't developed on, we ran five tasks from domains the benchmark doesn't cover, with the measures and pass criteria fixed in advance: the topic and the sentiment of financial news posts, the field of arXiv abstracts, the study type of PubMed abstracts, and the type of GitHub issues, the last three all published in 2026. Two of the questions, topic and sentiment, also occur in the benchmark. Each task was run once, and nothing was tuned on it.
Every pre-registered criterion passed. Without calibration, the model was even more overconfident than on the known tasks, with 95.3% stated confidence at 78.8% accuracy. The default temperature reduced the excess ECE from 0.151 to 0.047 and, measured as above, got 78% of the way in ECE to a per-task temperature, against 71% on the known tasks. With 100 labels, calibrate() raised accuracy by 2.7 percentage points, with gains on every task. Its 90% prediction sets contained the correct answer 91.2% of the time with 1.5 options on average, and only 0.8% of the samples exceeded the 10% error bound.
Two results are weaker. Without labels, the arXiv and GitHub tasks were still 0.08 away from calibrated. Both are harder than the model's confidence suggests, and their own best temperatures (2.8 and 3.5) are well above the default. With 100 labels, calibrate() brings both to 0.016 or below. Automation under the error bound was also lower: 9% of the answers at a 10% bound with 100 labels, compared with 23% on the known tasks. The test stops at the first threshold that fails, and on these tasks some of the most confident answers were already wrong, several apparently because of label errors. The error bound itself held. The full report has the numbers per task.
The chart shows how much of the calibration error each method leaves, and the list what each one needs and gives you, on the benchmark's 28 text datasets:
calibrate() row fits a temperature and a bias per option on 100 random labels per task, averaged over 20 draws.temperature="sentiment"): slightly better, if your task fits the family.calibrate() (about 100 labels): a temperature and a bias per option for your task; nearly calibrated probabilities, and 2.0 percentage points more answers right.predict_set() (same labels): a short list of options that contains the right one 90% of the time, 1.7 options on average.automate() (same labels): which answers to act on without review, with their error rate bounded; 23% of answers at a 10% bound.In practice: start with the default, which needs nothing from you. Once a decision runs at volume, label about 100 random inputs and let calibrate() tune the model to it.
We also tried several other published methods, on the same examples and under the same rules. None became a default; the calibration report has the details.
N/A. It made 24 of 28 datasets worse, because that answer is mostly real knowledge: for an empty input, neutral or other often is correct.permutations=4).option_mass.In the previous post's prompt, the state came first, followed by the question and the options. LLMs, however, read left to right: each position can only attend to what came before it. While reading the state, the model therefore doesn't yet know what it will be asked about it. The fix is to state the question and options before the state as well:
On the test halves of all 29 datasets, with two runs of each prompt, this raised mean accuracy by 1.6 percentage points, from 77.6% to 79.3%. It was better on 15 datasets, within a point on 12, and worse on 2 (Wilcoxon p = 0.002). On the 28 datasets that Jev answers, where the previous post found the two on par, GLM-5.3-Flash is now ahead: 79.8% against Jev's 77.5%, better on 16 datasets, tied on 8, and worse on 3.
The gain is concentrated where the model was unsure: +9.7 points in the least confident fifth of answers, and roughly zero elsewhere. The largest gains were on sentiment and intent tasks, such as +10.3 points on sst5 and +6.3 on toxic_conversations. It doesn't help with knowledge: on MMLU-Pro, accuracy went slightly down, from 63.1% to 61.9%. And it doesn't come from leaning more on the option names, since renaming every option to a synonym costs both prompts about the same (−7.7 and −8.1 points).
The default temperature still fits the new prompt. Refitting the formula on each prompt's answers, again leaving the scored dataset out, gives an excess ECE of 0.026 for the new prompt and 0.027 for the old one, so the new prompt is calibrated as well as the old one and needs no new constants. The shipped formula, fitted on the old prompt's answers, gives 0.016 on the new prompt and 0.023 on the old one. These numbers come from a different split of the test halves, so they don't compare directly with the 0.032 above.
The price is the question's tokens appearing twice: on average 1.6 times the input tokens, and therefore 1.6 times the cost. Up to a few dozen options, latency doesn't change measurably. With long option lists, it does: +110 ms at 77 options and +330 ms at 151.
When a request has several questions about the same state, the library asks the model once per question, and its optimize parameter chooses between two layouts. By default (optimize="accuracy"), each model call leads with its own question. With optimize="cost", every model call leads with all of the request's questions, so the calls share a prefix that the server can cache. With five questions per state, on a development set, the default raised accuracy by 2.2 percentage points and the shared layout by 0.9. The cache pays off only for long states, because Privatemode caches prefixes from about 2,300 tokens: on 2,000-token states, "cost" was the faster layout (670 ms median latency with 59% of the prompt cached, against 935 ms), while on short states it was the slower one (530 ms against 409 ms).
We also tried giving the model more to work with in the same forward pass, with filler tokens or more repetition, and nothing beat this prompt; the accuracy report has the details. Real thinking before the answer does more, but it turns milliseconds into seconds and is no longer a one-token decision.
JevBench is an independent benchmark for Jev-like decision models, with a public leaderboard. Our system isn't on the leaderboard yet, as submissions wait in a long queue. Until then, we answered JevBench's 231 public items the same way every other system did, without tuning anything on them, and compare accuracy on those.
With one token, GLM-5.3-Flash reaches 89.4% on them (the mean of two runs), up from 88.5% with the previous post's prompt. Jev 1.13.0 reaches 86.6% on the leaderboard. Ranked by public accuracy among the decision models on the leaderboard, as of 27 September 2026, our setup would place fifth of 90, behind one Gemma-based service at 92.8% and three systems at 89.6%, one item ahead. The leaderboard also lists general-purpose LLMs as baselines; those that reason before answering reach up to 99.6%, at seconds per decision, and are left out of the figure.
This comparison comes with caveats. Public items can be trained on or selected against, which is why JevBench also scores held-out items. On those, every system on the leaderboard is far less accurate (Jev: 36.7%), so public accuracy doesn't predict the official rank. JevBench's items are also closer to rules and judgment calls than our classification datasets. On MMLU-Pro, a knowledge and reasoning test, the picture reverses: Jev reaches a published 82.9%, and GLM-5.3-Flash in a single pass 61.9%. For decisions that require knowledge or several reasoning steps, a single token from a Flash-class model is not enough.
The playground in the previous post now uses everything described here: the question-first prompt and the default temperature for GLM-5.3-Flash. Each answer shows the temperature its probabilities were scaled by. The trick question there remains a good reminder of what calibration can't do: the model answers it wrong in one token, and calibration only makes it less sure of the wrong answer.
One more change since the previous post is easy to miss: the prefill is now answer= instead of choice_index:. After a colon, the model would rather write a space than a digit, and the chat template strips trailing spaces, so less than 0.1% of the probability landed on the options. With answer=, nearly all of it does. Accuracy stayed the same, and all numbers in this post use answer=.
answer=, it puts about a third of its probability elsewhere, mostly on a space.Both the library and the benchmark are open source. The benchmark repository contains the full calibration report in three parts, with every table in this post and many more, the code that produced them, the raw calibration runs, and tests that check the library's calibration code against the report. The library's README explains which guarantee holds when, and its AGENTS.md tells coding agents what an implementation has to get right.
edgelesssys/privatemode-decisionsThe Python library: token oracle, prompt, masking and renormalization, against any vLLM-backed endpoint. edgelesssys/privatemode-decisions-benchmarkThe benchmark: methodology, frozen dataset specs, harness and aggregation. Every number in this post can be recomputed from it.With Privatemode, these decisions are protected by confidential computing: your data stays encrypted in memory even during processing, and the client verifies the deployment's attestation report before sending anything. You can read more on our security page.
Create an account, then give this post to Claude Code, Codex, or the coding tool of your choice. Together with the GitHub repository, it has everything needed to build typed decisions into your own code, protected by confidential computing.
OpenAI now says AI inference belongs in confidential computing. What the DevDay announcement contains, what it leaves open, and three questions that separate a verifiable property from a promise.
Typed decisions in a single forward pass, about as accurate as Jev and faster from Europe. With a playground and a benchmark on 29 datasets.
