Research / Paperworking draft
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noise
Abstract
Typed-decision models answer in fixed-shape fields, such as a yes or no, one option from a list, or a score, in a single forward pass, and they are sold as unable to hallucinate. I tested what that guarantee leaves out. BEMA is an open, from-scratch reproduction of the interface: one small transformer encoder with binary, 77-way and regression heads, trained only on human-labeled data and calibrated with temperature, Platt, isotonic and split-conformal methods. On clean held-out text it looks trustworthy: temperature scaling cuts the binary head’s expected calibration error from 0.0159 to 0.0046. Then I added two adjacent-character typos per query, the ordinary fast-typing kind. Accuracy on the 77-way head fell from 81.9% to 51.2% (n = 3,080), while mean confidence fell only from 0.791 to 0.612. The same gap appears on a second, 151-way dataset. Retraining on typo-augmented data halves the accuracy drop, and the gain carries over to typo types it never trained on, but it does not close the gap. The output guarantee is real: the model cannot emit an answer outside its schema. It can still be confidently wrong on input squarely inside its own domain. A noise-robustness number belongs next to ECE whenever a model’s confidence is sold as a safety signal.
Key findings
- Two adjacent-character typos per query cut the 77-way head’s accuracy from 81.9% to 51.2% (n = 3,080), while mean confidence fell only from 0.791 to 0.612.
mit37/BEMA/blob/d68f1a4/FINDINGS.md#L164-L173 - The gap reproduced on CLINC150 with a separate model and tokenizer: accuracy 71.80% → 45.04%, confidence 0.7514 → 0.5615 (n = 5,500).
mit37/BEMA/blob/d68f1a4/CLINC150.md#L46-L49 - Typo-augmented retraining cut the accuracy drop from 31.4 to 16.0 points at no cost to clean accuracy, and was 11 to 15 points more accurate on typo types it never trained on, but did not close the gap.
mit37/BEMA/blob/d68f1a4/FINDINGS.md#L460-L526 - On clean text, no calibration method won everywhere: temperature scaling cut the binary head’s ECE from 0.0159 to 0.0046 but raised the 77-way head’s from 0.0210 to 0.0371; isotonic regression brought it to 0.0113.
mit37/BEMA/blob/d68f1a4/results_log.csv#L35-L41 - Split conformal prediction met its 90% coverage target on all three heads, within finite-sample noise, but valid is not useful: the 77-way sets average 7.07 of 77 classes.
mit37/BEMA/blob/d68f1a4/CONFORMAL.md#L49-L61
Introduction
On 15 September 2026, TypeSafe AI introduced Jev and a category it calls System One models: models that return typed decisions instead of text (TypeSafe, 2026). A caller sends a state and a schema of questions. Each answer comes back as an option chosen from a list, a score on a rubric, or a 0-to-1 value for whether a statement is true, and every question is evaluated in parallel against the same state (TypeSafe docs). The launch post makes two claims that matter for anyone who wants to put such a model in front of money or people. First, it plots Jev’s hallucination rate at 0% because every output matches the schema, and its fine print is candid about the basis: “Our number is not empirical.” Second, it describes the model as “Calibrated: higher confidence means higher accuracy.”
Those are two different kinds of claim. The first is about form. A model whose output layer is a fixed-shape projection over a declared set of answers cannot emit a malformed or off-schema answer, by construction. That is true, and autoregressive language models, which can emit any token sequence, do not share it. The second is about truth. It says the reported confidence tracks how often the answer is right, and that is an empirical property. It holds only as far as the data it was measured on.
I wanted to know whether the second claim survives realistic input. I could not test Jev itself, so I built an open reproduction of its interface shape: typed state in, one forward pass, a typed decision and a calibrated confidence out. BEMA is not Jev. It shares none of Jev’s weights, training data or scale, and nothing below measures Jev’s accuracy. What it can test is whether the safety story attached to this interface holds in a model built the same way.
This paper reports five results:
- An open reproduction. One shared encoder answers three typed questions from one pooled state. It is trained from scratch on human-labeled data only, under a per-dataset license record, and released under MIT.
- No calibration method wins everywhere. Temperature scaling was the best simple fix for the binary head and made the 77-way head worse.
- The headline. Two ordinary typos per query nearly halve the 77-way head’s accuracy while its confidence falls by less than a quarter. The pattern reproduces on a second dataset with a separate model and tokenizer.
- A partial fix. Typo-augmented training halves the damage, and the gain carries over to typo types it never saw, but it does not remove the gap, and I name its confounds.
- Valid is not useful. Split conformal prediction delivers its coverage guarantee on every head, within finite-sample noise, but the 77-way sets are wide, and adaptive intervals for the regression head cover one-star reviews only about half the time.
How this study was run. I specified BEMA as a research plan, a baseline build followed by five follow-up workstreams, and directed Claude Code agents to implement it, run it and write up the results. 47 of the repository’s 49 commits are agent commits; the other two are my merges of the agents’ pull requests. The scope calls came back to me: when the agents flagged a dataset’s licensing as needing a human decision, I dropped it. The plan also carried one standing rule, not to fake or approximate an experiment the environment could not run, and two experiments below are reported as not run for exactly that reason. Every number in this paper is copied from the committed files in mit37/BEMA, and every table cites the file and line each value came from.
Model: one encoder, three typed heads
The model is a small transformer encoder (Vaswani et al., 2017) with three decision heads that read the same state vector.
- Tokenizer. A byte-level BPE vocabulary of 8,000 tokens (Sennrich et al., 2015), trained only on the union of the three training splits, with inputs capped at 64 tokens.
- Trunk. Sinusoidal positions, then three encoder layers with a model width of 96, 6 attention heads, a feed-forward width of 384 and dropout of 0.15. Masked mean pooling and a LayerNorm reduce the sequence to one state vector.
- Heads. Three two-layer MLPs read that vector. Noul gives one logit for a yes-or-no decision (here, spam or not). Choice gives logits over a fixed label set (here, 77 banking intents). Score gives one logit whose sigmoid is a continuous score in [0, 1] (here, a normalized star rating).
forward_all()runs the trunk once and reads every head from it.
The output guarantee is structural. Each head is a fixed-shape projection, so the model cannot return anything but a value its schema declares. Nothing about calibration or accuracy follows from that.
Training draws one batch per task per step and sums three losses: binary cross-entropy with a positive-class weight for the spam imbalance, cross-entropy for Choice and mean squared error for Score. The optimizer is AdamW (learning rate 3 × 10⁻⁴, weight decay 10⁻⁴) with gradient-norm clipping at 1.0, for 12 epochs, keeping the checkpoint with the best combined validation score.
Two constraints shaped the model. First, no pretrained weights: the build environment blocked model hubs, so every number here comes from an encoder trained from scratch on a few thousand examples per task. That is the largest single limit on quality in this study. Second, scaling up is not free even at this size. The first attempt to double the encoder at the original learning rate, without clipping, diverged: validation accuracy collapsed from about 90% to 15% partway through training. A lower learning rate and gradient clipping fixed it.
Data and licensing
Every label in this study was written by a person. None was generated by a language model. Each dataset’s license was checked before use, and the repository keeps one record of what was verified and how (DATA_LICENSES.md).
| Head | Dataset | Size used | License, and how it was checked |
|---|---|---|---|
| Noul | SMS Spam Collection (Almeida et al., 2011) | 5,572 messages; 836 in test | Not independently verified; treated as research use only |
| Choice | BANKING77 (Casanueva et al., 2020) | 10,003 train, 3,080 test; 77 intents | CC BY 4.0, read from the first-party LICENSE file |
| Score | Amazon Fine Food Reviews (McAuley & Leskovec, 2013) | first 10,112 rows; 1,517 in test | CC0, corroborated by two independent listings (a weaker check) |
| Choice, second domain | CLINC150 (Larson et al., 2019) | 15,000 / 3,000 / 4,500 in-scope, plus 100 / 100 / 1,000 out-of-scope | CC BY 3.0, read from the first-party LICENSE file |
Two datasets were dropped on licensing grounds, and both decisions cost results. The Score head first used the STS Benchmark. A review found that its text mixed several sub-sources with unresolved, non-uniform terms, so it was replaced outright with Amazon Fine Food Reviews rather than worked around. Later, a toxicity dataset considered for a second domain carried a permissive top-level license alongside separate academic-use requests from its authors. It was rejected as ambiguous, and that workstream stopped at one new dataset instead of reaching for a less scrutinized one.
Each dataset has its own held-out test split, and calibration parameters were fit on validation data only. Nothing that touched a fitting step was used to measure it.
Calibration methods
Expected calibration error (ECE) sorts predictions into bins by confidence and averages the gap between accuracy and confidence in each bin, weighted by the bin’s share of the data. This study uses top-label ECE with ten equal-width bins:
Three post-hoc methods were fit on the validation split and scored on test:
- Temperature scaling divides the logits by one scalar fit by minimizing validation loss (Guo et al., 2017). It cannot change which answer is predicted.
- Platt scaling fits a logistic map on the binary head’s logit.
- Isotonic regression, for the 77-way head, fits a monotone map from the top softmax probability to the chance of being right. The predicted class always comes from the raw softmax; only the reported confidence is recalibrated. An earlier version of this path overwrote a probability in place and then re-took the argmax, which could silently change the prediction. An agent-run review of the repository caught it, and it was fixed before any number reported here.
Calibration scores average over a test set. Split conformal prediction makes a different promise: a set of answers, or an interval, that contains the truth with probability at least 1 − α on exchangeable data, with no assumption about the model (Angelopoulos & Bates, 2021). The threshold is the finite-sample quantile of nonconformity scores on held-out calibration data:
The Noul head uses the least-ambiguous threshold rule (Sadinle et al., 2016), the Choice head uses non-randomized adaptive prediction sets (Romano et al., 2020) and the Score head uses absolute residuals. To keep the conformal threshold independent of the temperature fit, the validation split is halved: the temperature is refit on one half and the conformal scores come from the other.
Clean results
On clean held-out text, the model behaves the way the pitch says it should.
| Head | Dataset | Output | Test n | Accuracy / MAE | ECE, raw | ECE, temperature | Best ECE (method) |
|---|---|---|---|---|---|---|---|
| Noul (binary) | SMS Spam Collection | spam / ham | 836 | 97.85% | 0.0159 | 0.0046 | 0.0030 (Platt) |
| Choice (77-way) | BANKING77 | 1 of 77 intents | 3,080 | 81.88% | 0.0210 | 0.0371 | 0.0113 (isotonic) |
| Score (regression) | Amazon Fine Food Reviews | rating in [0, 1] | 1,517 | MAE 0.1872 · r 0.534 | n/a | n/a | n/a |
Held-out test splits the calibration step never touched. Accuracy is unchanged by temperature scaling; ECE is top-label with 10 equal-width bins. The Score head is a regression and is not calibrated the same way.
The binary head is accurate and well calibrated after scaling. The 77-way head reaches 81.88% over 77 real classes from a small encoder trained from scratch. The Score head learns real signal: a Pearson correlation of 0.534 with held-out star ratings, and a mean absolute error of 0.1872 on a 0-to-1 scale. That is a modest result, and the paper claims no more for it.
The method comparison is less tidy.
| Head | Raw | Temperature | Platt | Isotonic |
|---|---|---|---|---|
| Noul ECE (n = 836) | 0.0159 | 0.0046 | 0.0030 | 0.0049 |
| Choice ECE (n = 3,080) | 0.0210 | 0.0371 | n/a (binary only) | 0.0113 |
Test-set ECE after fitting each method on the validation split only. Lower is better. Temperature scaling was the worst option for the 77-way head on this run.
Platt scaling beat temperature scaling on the binary head. On the 77-way head, temperature scaling made ECE worse than doing nothing, 0.0210 to 0.0371, and isotonic regression was the only method that helped. This is specific to this run: on earlier runs of the same pipeline, temperature scaling helped the 77-way head. The practical lesson is the same either way. A calibration step should try more than one method per head and check each on held-out data, because a single global temperature is not guaranteed to help (Kull et al., 2019).
How much of that is noise? For the binary head, three independent seeds, each with its own split, tokenizer and initialization, gave accuracy of 97.85% to 98.33% (mean 98.01%, standard deviation 0.28 points). Calibrated ECE ranged from 0.0039 to 0.0070 (mean 0.0054, standard deviation 0.0015), and temperature scaling cut ECE by 64.5% to 74.7% on every seed. A second sweep retrained the multitask model with the same three seeds on one fixed split. The 77-way head’s accuracy was 81.95% ± 0.11 points, and the Score head’s Pearson r was 0.561 ± 0.023, so the single-run r of 0.534 above is the lowest of the three. That sweep varies only the training seed, so split-to-split variance for those two heads is still unmeasured.
Split conformal prediction delivered its guarantee on every head, within finite-sample noise (Score’s 89.85% is a hair under the 90% target at n = 1,517), and showed how far “valid” can sit from “useful”.
| Head | Method | Calib / test n | Coverage | Set size / interval width | Read |
|---|---|---|---|---|---|
| Noul | Threshold (LAC) | 418 / 836 | 90.91% | avg 0.92 of 2 | 8.37% empty sets |
| Choice (BANKING77) | APS (non-randomized) | 750 / 3,080 | 98.83% | avg 7.07 of 77 | valid, conservative, wide |
| Score, constant width | Absolute residual | 759 / 1,517 | 89.85% | mean 0.656 on [0, 1] (raw 1.066) | valid, wide: about 2.6 stars |
| Score, CQR | CQR | 759 / 1,517 | 90.11% | mean 0.608 on [0, 1] | 1-star reviews covered 52.5% |
| Choice (CLINC150) | APS (non-randomized) | — / 5,500 | 98.27% | avg 17.31 of 151 | valid, conservative, wide |
Target coverage 90% (alpha = 0.10). Calibration scores come from half of the validation split, disjoint from the temperature fit and from the test set. Score intervals are clipped to the valid [0, 1] range.
The binary sets are tight, and when the method commits to one class it is right 99.2% of the time. The cost is that 8.37% of test messages get an empty set, an honest refusal that a typed interface has no field for. The 77-way head needs an average of 7.07 of its 77 classes to reach coverage, which is informative but not a shortlist. The Score interval is valid but wide. Its raw constant width is 1.066, but the label can only lie in [0, 1], so clipping loses no coverage; clipped, it averages 0.656, about 2.6 of the four steps on the star scale, and no interval is narrower than two stars. The cause is a right-skewed error distribution (median absolute error 0.090, 90th percentile 0.538): 61% of test reviews are five-star, which the model predicts well, and one width has to cover the hard tail. Conformalized quantile regression (Romano et al., 2019) adapts the width to the input. It kept coverage at 90.11% with a mean width of 0.608, and 41% of its intervals span less than two stars. Its coverage by true rating shows what a marginal guarantee hides: 97.0% for five-star reviews, 52.5% for one-star reviews. The 90% is an average over a test set dominated by the easy case. A deployment that cares about negative reviews has to measure that group directly.
The first crack in the safety story shows up before any typos. Fed real SMS messages, text with no valid banking intent at all, the 77-way head does not refuse. Its mean confidence falls from 0.7906 to 0.4097, far above the 1/77 ≈ 0.013 of a model that knows it has no answer. The schema guarantees the answer is one of 77 banking intents. It does not guarantee that any of them applies.
One architecture change was also tested, across the same three training seeds. Swapping mean pooling for a learned attention pool made no measurable difference: every paired difference is smaller than its own seed-to-seed spread, and the edge a first single-seed run showed for attention pooling did not hold up.
| Measure | Mean pool | Attention pool | Paired difference |
|---|---|---|---|
| Noul accuracy (n = 836) | 0.9801 ± 0.0050 | 0.9809 ± 0.0041 | +0.0008 ± 0.0014 |
| Choice accuracy (n = 3,080) | 0.8195 ± 0.0011 | 0.8202 ± 0.0143 | +0.0008 ± 0.0136 |
| Choice ECE, calibrated | 0.0341 ± 0.0042 | 0.0314 ± 0.0046 | −0.0027 ± 0.0050 |
| Score MAE (n = 1,517; lower is better) | 0.1858 ± 0.0055 | 0.1809 ± 0.0041 | −0.0049 ± 0.0095 |
| Score Pearson r | 0.5605 ± 0.0234 | 0.5666 ± 0.0199 | +0.0061 ± 0.0190 |
Mean ± sample standard deviation over seeds 42, 123 and 2024, same data split, everything but the pooling step held fixed. Every paired difference is smaller than its own seed-to-seed spread.
Robustness under typos
The out-of-scope result needs a change of domain. This one does not.
I took every query in BANKING77’s test split, 3,080 real customer questions, and applied two random adjacent-character swaps to each. It is the error a person makes typing fast: “How do I locate my card?” becomes “Howd o I locate my crad?”, which any human reads without trouble. Same schema, same task, same domain.
Accuracy fell from 81.9% to 51.2%. Mean confidence fell from 0.791 to 0.612. In relative terms, accuracy dropped by 37.5% and confidence by 22.6%. On clean text the head was slightly underconfident, with mean confidence 2.8 points below accuracy. With two typos per query it was overconfident by 10.0 points.
Each point is one model on one test set: mean confidence across the set (x) against accuracy (y). On the diagonal, confidence matches accuracy. Below it, the model is overconfident. A line joins each model’s clean point (hollow) to its point with two typos per query (filled). This is not a binned reliability diagram. Per-example predictions were never persisted (checkpoints and pickles are gitignored), so per-bin accuracy is unavailable, on clean or perturbed text. Each point is a whole-test-set average, which can hide miscalibration that cancels across bins.
| Model and test set | Test n | Accuracy, clean | Accuracy, typos | Confidence, clean | Confidence, typos | Accuracy drop (relative) | Confidence drop (relative) | Overconfidence under typos |
|---|---|---|---|---|---|---|---|---|
| BANKING77, baseline (first draw) | 3,080 | 81.9% | 51.2% | 0.791 | 0.612 | 37.5% | 22.6% | +10.0 pts |
| BANKING77, baseline (paired draw) | 3,080 | 81.88% | 50.52% | 0.7906 | 0.6122 | 38.3% | 22.6% | +10.7 pts |
| BANKING77, typo-augmented retrain | 3,080 | 83.47% | 67.44% | 0.8133 | 0.7046 | 19.2% | 13.4% | +3.0 pts |
| CLINC150, standalone 151-way model | 5,500 | 71.80% | 45.04% | 0.7514 | 0.5615 | 37.3% | 25.3% | +11.1 pts |
Two random adjacent-character swaps per query ("How do I locate my card?" becomes "Howd o I locate my crad?"), applied to every test query. Confidence is the temperature-scaled top-label probability. Overconfidence under typos is mean confidence minus accuracy.
Nothing else in the study would have caught this. Every ECE figure above was measured on clean, unperturbed text, where the 77-way head’s calibration error was 0.021 raw and 0.037 after temperature scaling. On the noise every text box receives, its confidence stopped tracking its accuracy. Earlier work shows that character-level noise breaks neural text models (Belinkov & Bisk, 2017; Pruthi et al., 2019) and that calibration degrades under dataset shift (Ovadia et al., 2019). The point here is narrower and more practical: for a typed-decision model sold on its confidence, ordinary typing noise is enough to break the confidence–accuracy link, with no adversary and no change of domain.
What the tokenizer does, and which claim survives it. The collapse runs through a specific mechanism. Byte-level BPE never produces an unknown token, but a swap often splits one clean token into several weak pieces: “available” becomes “vaailable”, which tokenizes as v, aa, ila, ble instead of one available token. Four of five sampled queries fragmented by two to five extra tokens. This tokenizer is small (8,000 tokens) and trained on a narrow corpus, so it is probably more fragile than a pretrained tokenizer with a vocabulary around 30,000. So the finding splits in two.
- The calibration gap is the claim this paper carries forward. Whatever corrupts the representation, a calibrated model’s confidence should fall as far as its accuracy does. Here it did not.
- The size of the collapse is not claimed for production-grade encoders. A 30.7-point drop is this implementation’s number.
The experiment that would separate the two is to put a pretrained subword encoder under the Choice head, retrain and rerun the typo test. It could not run: the build environment refused connections to the model hubs, which was re-checked immediately before the write-up. Under the no-approximation rule, it is reported as unresolved. Anyone with hub access can run it against the repository as it stands.
Three further caveats travel with these numbers. The confidences come from the temperature-scaled head, which on clean data was the worse-calibrated of the three options, and isotonic recalibration was never applied under typos. The headline run measured only mean confidence against accuracy. A later held-out-noise test reports ECE under each corruption (next section), but per-example predictions were never saved, so there is no binned reliability diagram for perturbed text; FIG 5 plots test-set averages for that reason. And the typo draw is random. Three draws of two swaps on the same checkpoint gave 51.2%, 50.52% and 50.00% accuracy, but perturbation-seed variance was not measured beyond those draws.
Mitigation: training on typos
If the model has never seen a typo, show it some. The augmentation run duplicated half of each task’s training rows with the same two-swap perturbation, taking the 77-way head from 8,503 to 12,754 training rows. It then retrained a fresh encoder with the same architecture and recipe and recalibrated it. Validation and test data were left untouched. Both models were then evaluated on the identical perturbed text, and the baseline was re-scored live from its checkpoint, not copied from earlier results.
It helped, and it did not finish the job. Typo-set accuracy rose from 50.52% to 67.44%, so the drop under typos shrank from 31.4 points to 16.0. Clean accuracy rose too, from 81.88% to 83.47%. Mean confidence under typos went from 10.7 points above accuracy to 3.0 points above it. Typo-set accuracy is still 16 points below clean accuracy.
The augmented model only ever saw adjacent swaps, so its gain could have been memorization of that one corruption. It was not. On the same 3,080 queries under four corruptions absent from training (keyboard-neighbor substitutions, deletions, insertions, and a mix of all three), it was 11 to 15 points more accurate than the baseline, and its ECE fell by roughly half or more. It is still not a cure: accuracy under that noise stays 13 to 24 points below clean accuracy, the model remains overconfident under every corruption, and none of these synthetic corruptions stands in for real user-typed text.
Two confounds limit what this shows:
- The augmented run saw about 50% more optimizer steps per epoch: 399 against the baseline’s 266, because steps scale with training rows.
- A control that duplicates rows without typos was not run, so part of the gain may be more data rather than typo robustness.
Generalization: a second dataset
A result on one dataset could be an artifact of that dataset or its tokenizer. So the same perturbation was run on CLINC150: 150 crowd-sourced intents across many domains plus an explicit out-of-scope class, 151 classes in all. It used a standalone model with the same architecture and recipe, and its own 4,000-token tokenizer trained on that corpus alone.
The pattern held. Accuracy fell from 71.80% to 45.04% while mean confidence fell from 0.7514 to 0.5615, relative drops of 37.3% and 25.3%, close to BANKING77’s. Different data, domains, vocabulary and class count gave the same gap. On clean text, temperature scaling cut this model’s ECE from 0.1375 to 0.0412, which again says nothing about what happens under noise.
CLINC150 also tests something BANKING77 cannot: what a model does when “none of the above” is a legal answer. As a label, it barely gets used. The model picked it for 16.1% of true out-of-scope queries, against 84.2% accuracy on in-scope queries. That is not class imbalance: CLINC150’s training set has exactly 100 examples for every class, out-of-scope included. The difference is what those 100 examples must cover. An in-scope intent is one narrow request; out-of-scope is everything else. The confidence score carries more of the signal than the label does. Top-class probability averages 0.82 on in-scope queries and 0.44 on out-of-scope ones (AUROC 0.88). Answering “out of scope” whenever it falls below a threshold tuned on validation data raises out-of-scope recall to 85.7%, at a price: in-scope accuracy falls from 84.2% to 74.5%, and more than half of the “out of scope” answers are then wrong. So the model partly knows when it does not know, as a tunable trade-off rather than something the schema gives for free. The overall 71.80% blends two very different numbers, so it is not a single meaningful accuracy.
Speed and cost against language models: not measured
Speed is part of the typed-decision pitch: TypeSafe reports end-to-end responses 40 to 200 times faster than frontier models on queries shaped for its model (TypeSafe, 2026). BEMA’s own side of that ratio is measured: about 1.4 ms mean latency on a CPU for all three heads in one forward pass. That is one fixed sentence at batch size 1, 200 timed trials after 10 warm-up runs, with tokenization excluded.
The other side was never measured. The build environment had no API credentials and no local model server, so no language model answered the same questions under the same clock. This paper therefore makes no speed or cost comparison. A small forward pass beating an API round trip is plausible; the ratio is untested here.
Limitations
- Scale. A 96-wide, three-layer encoder trained from scratch, with no pretrained weights. Every accuracy figure is a toy-scale figure.
- Statistics. One held-out split per dataset, test sets of 836 to 5,500 examples and no bootstrap intervals. Seed variance covers fresh splits for the binary head, but only training seeds on one split for the 77-way and Score heads.
- The perturbation. The headline uses one noise type (adjacent swaps) at two per query; three random draws of it on one checkpoint agree within 1.2 points. The held-out-noise test adds three more types and a heavier swap count, all synthetic; none comes from real users.
- What was measured under noise. Test-set mean confidence against accuracy for the headline, and per-corruption ECE in the held-out-noise test, but no binned reliability diagram.
- The mitigation. Confounded by extra optimizer steps and extra data, and tested on synthetic noise rather than real user typos.
- Data. The SMS Spam Collection’s license is not independently verified.
- Scope. BEMA reproduces an interface, not a product. Nothing here measures Jev, and the speed and cost comparison was not run.
Conclusion
“Cannot hallucinate” is a claim about form, and in a typed-decision model it is true. The claim that matters when the decision moves money or reaches a person is different: that the model knows when it does not know. In this reproduction, textbook calibration made the clean-test numbers look excellent. Then two typos per query nearly halved accuracy while confidence fell by less than a quarter, on a second dataset as well as the first. Training on typos narrowed the gap without closing it.
Anyone evaluating a typed-decision model, or selling one, should report four things next to clean-test ECE. First, a robustness number under realistic noise. Second, out-of-scope behavior. Third, calibrated sets or intervals, where a single answer is not safe. Fourth, whether the calibration method was chosen per head on held-out data. The next experiments are already specified: rerun with per-example outputs saved, so reliability diagrams can be drawn under noise; vary typo seeds and test on real user-typed text; put a pretrained encoder under the Choice head; run the language-model speed and cost comparison with a real API key; and conformalize the Score intervals within rating groups, so one-star reviews get the coverage the average promises.
Reproduce
Every number above comes from these scripts, on CPU, in the repository at commit d68f1a4:
pip install -r requirements.txt
python3 prepare_multitask.py && python3 train_multitask.py && python3 calibrate_multitask.py
python3 calibration_sweep.py # FIG 1–2
python3 ood_stress_test.py # out-of-scope confidence
python3 adversarial_stress_test.py # typo attack and tokenizer audit
python3 conformal.py # FIG 3
python3 score_cqr.py # adaptive Score intervals
python3 augment_and_retrain.py # typo-augmented retrain
python3 heldout_noise_test.py # typo types it never trained on
python3 clinc150_pipeline.py # second dataset
python3 clinc150_oos_threshold.py # out-of-scope rejection
python3 seed_variation_sweep.py # three-seed sweep, binary head
python3 multiseed_sweep.py # FIG 4, three training seeds
python3 benchmark_speed.py # latency
References
- Almeida, T. A., Gómez Hidalgo, J. M., and Yamakami, A. (2011). Contributions to the study of SMS spam filtering: new collection and results. DocEng ’11. doi:10.1145/2034691.2034742
- Angelopoulos, A. N., and Bates, S. (2021). A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification. arXiv:2107.07511
- Belinkov, Y., and Bisk, Y. (2017). Synthetic and Natural Noise Both Break Neural Machine Translation. arXiv:1711.02173
- Casanueva, I., Temčinas, T., Gerz, D., Henderson, M., and Vulić, I. (2020). Efficient Intent Detection with Dual Sentence Encoders. arXiv:2003.04807
- Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017). On Calibration of Modern Neural Networks. arXiv:1706.04599
- Kull, M., Perello-Nieto, M., Kängsepp, M., Silva Filho, T., Song, H., and Flach, P. (2019). Beyond temperature scaling: Obtaining well-calibrated multiclass probabilities with Dirichlet calibration. arXiv:1910.12656
- Larson, S., Mahendran, A., Peper, J. J., Clarke, C., Lee, A., Hill, P., Kummerfeld, J. K., Leach, K., Laurenzano, M. A., Tang, L., and Mars, J. (2019). An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction. arXiv:1909.02027
- McAuley, J., and Leskovec, J. (2013). From Amateurs to Connoisseurs: Modeling the Evolution of User Expertise through Online Reviews. arXiv:1303.4402
- Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J. V., Lakshminarayanan, B., and Snoek, J. (2019). Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift. arXiv:1906.02530
- Pruthi, D., Dhingra, B., and Lipton, Z. C. (2019). Combating Adversarial Misspellings with Robust Word Recognition. arXiv:1905.11268
- Romano, Y., Patterson, E., and Candès, E. J. (2019). Conformalized Quantile Regression. arXiv:1905.03222
- Romano, Y., Sesia, M., and Candès, E. J. (2020). Classification with Valid and Adaptive Coverage. arXiv:2006.02544
- Sadinle, M., Lei, J., and Wasserman, L. (2016). Least Ambiguous Set-Valued Classifiers with Bounded Error Levels. arXiv:1609.00451
- Sennrich, R., Haddow, B., and Birch, A. (2015). Neural Machine Translation of Rare Words with Subword Units. arXiv:1508.07909
- TypeSafe AI (2026). Introducing System One Models & Jev. typesafe.ai; documentation at docs.typesafe.ai.
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention Is All You Need. arXiv:1706.03762