Work / Research
BEMA
Stress-tested a “can’t-hallucinate” model design: 2 typos cut accuracy 82% → 51%, confidence only 0.79 → 0.61.
- Python
- PyTorch
- Hugging Face tokenizers (byte-level BPE, trained from scratch)
- scikit-learn
- NumPy
Live demo
Typo attack
The perturbation BEMA used, running live: two adjacent-character swaps. Then the measured cost of those two swaps, from the repo.
CleanHow do I locate my card?
What the model sawHowd o I locate my crad?
Two adjacent-character swaps per query, exactly as BEMA’s stress test applies them. mit37/BEMA/adversarial_stress_test.py#L193-L200 (opens in a new tab)
Why two swaps hurt so much
available →availablevaailable →vaailable
One swap splits one familiar token into four fragments the small tokenizer has barely seen. mit37/BEMA/FINDINGS.md#L394-L398 (opens in a new tab)
Two typos cut accuracy by 37.5%, but mean confidence by only 22.6% (relative, BANKING77).
Clean vs. two typos , full test splits.
BANKING77 · 77 classes · n = 3,080
Under typos the model is 10.0 points more confident than it is accurate.
CLINC150 · 151 classes · n = 5,500
Under typos the model is 11.1 points more confident than it is accurate.
The fix, partly: train on typos
BANKING77 · same typo draws · n = 3,080
Under typos the model is 10.7 points more confident than it is accurate.
- Confidence under typos is mean top-label confidence from the temperature-scaled head, not a calibration error. A later held-out-noise run reports ECE per corruption, but no reliability diagram exists for typo inputs. mit37/BEMA/adversarial_stress_test.py#L203-L234 (opens in a new tab)
- The size of the accuracy drop is specific to this small, from-scratch 8,000-token tokenizer; it is not claimed for pretrained tokenizers. The gap between confidence and accuracy is the finding. mit37/BEMA/FINDINGS.md#L403-L422 (opens in a new tab)
- Three random typo draws on the same baseline gave 51.2%, 50.52% and 50.00%; variance across typo draws is not reported. mit37/BEMA/FINDINGS.md#L464-L519 (opens in a new tab)
No prediction is shown for your text: BEMA hasn’t exported per-example outputs yet, and this demo never guesses one. The typo is real; the numbers are the repo’s measured aggregates.
TL;DR
- Typed-decision models answer in fixed-shape fields (a yes or no, a pick from a list, a number) in one forward pass, and the design is pitched as unable to hallucinate. I built an open, MIT-licensed reproduction of that interface to test what the claim leaves out.
- On clean data it looks trustworthy. Under two ordinary typos, 77-way accuracy fell from 81.9% to 51.2% (n = 3,080) while mean confidence only fell from 0.79 to 0.61.
- Calibration measured on clean text never sees that gap. The result reproduced on a second domain, and I report the negative results next to the positive ones.
The problem
A model with no free text to generate can’t invent a citation. It can still be wrong, and the question that matters is whether it knows when it is. The standard reliability metric, expected calibration error (ECE), is measured on clean test sets. Real inputs have typos. Fintech and quant teams live on that difference: a model that is confidently wrong under ordinary noise is worse than one that is uncertain.
What I built
I specified the work as a five-workstream research plan and directed Claude Code agents through the build, the experiments and the write-up. The model is trained from scratch on real human-labeled data only, never on LLM-generated labels.
- Model. A small transformer encoder (3 layers, 6 heads, width 96) over a byte-level BPE tokenizer trained on the repo’s own training splits. Three typed heads read one shared state vector: Noul (binary: SMS spam), Choice (77-way banking intent: BANKING77) and Score (1–5 star ratings as regression). One forward pass answers all three.
- Calibration. Temperature, Platt and isotonic recalibration, measured as top-label ECE with 10 bins. Multiclass isotonic recalibrates only the reported confidence and never changes the prediction.
- Conformal prediction. Split conformal sets for all three heads, with the validation split halved so recalibration and conformal calibration never share data.
- Stress tests. Two random adjacent-character swaps on every BANKING77 test query, spam-evasion tricks (spacing, leetspeak, case), cross-domain inputs, and a tokenizer-fragmentation audit.
- Follow-ups. Typo-augmented retraining, tested on typo types it never saw; a standalone 151-class model on CLINC150, with out-of-scope rejection by confidence threshold; adaptive conformal intervals for Score; and an attention-pooling ablation across three training seeds.
Key decisions
- Decision: real human labels only. Why: a test of hallucination can’t rest on labels a model generated. Trade-off: smaller datasets, and each one needs its own licensing check.
- Decision: drop STS-B. Why: the agents flagged mixed licensing terms across its sub-sources, and I made the call to cut it; the repo keeps a per-dataset licensing record. Trade-off: one fewer regression benchmark.
- Decision: report the pretrained-encoder comparison as not run rather than approximate it. Why: the build environment blocked the model hubs, so every encoder is trained from scratch, and the plan’s standing rule was never to fake an experiment that couldn’t run. Trade-off: a small model, and the biggest open question stays open.
- Decision: report what didn’t work. Why: a calibration study that hides its failures isn’t one. Trade-off: the write-up reads less like a highlight reel, which is the point.
The hard part
Is the typo collapse a property of the architecture, or an artifact of my tokenizer? The 81.9% → 51.2% number is striking, and it would be easy to over-claim it.
The audit showed that byte-level BPE never produces an unknown token, but a two-character swap often splits one clean token into several weak ones: “available” becomes “vaailable”, which tokenizes as v, aa, ila, ble. Four of five sampled queries fragmented by two to five extra tokens. So I split the finding into two claims:
- Confidence doesn’t track accuracy under ordinary noise. A roughly 23% relative confidence drop against a roughly 37% relative accuracy drop. This is the headline, and it should hold whatever the tokenizer.
- The size of the collapse is explicitly not claimed for pretrained tokenizers with larger vocabularies. The experiment that would settle it, a pretrained encoder, couldn’t run in the build sandbox, so it stays listed as unresolved rather than approximated.
A second catch came from an agent-run review across correctness, statistics, licensing, reproducibility and security, with adversarial verification of each finding. It found that the multiclass isotonic path overwrote a probability in place and re-took the argmax, which could silently flip predictions. The fix derives predictions from the raw softmax and recalibrates only the reported confidence.
Results
- Clean data. Noul reaches 97.85% accuracy, and temperature scaling cuts its ECE from 0.0159 to 0.0046. Choice scores 81.88% on BANKING77’s full 3,080-query test set. Score reaches a test MAE of 0.187 with Pearson r = 0.53, the lowest of three training seeds (0.53 to 0.58).
- Under typos. Choice accuracy falls from 81.9% to 51.2% while mean confidence only falls from 0.79 to 0.61. On CLINC150, a separate 151-class model falls from 71.8% to 45.0% while confidence falls only from 0.75 to 0.56 (n = 5,500).
- Mitigation. Retraining on typo-augmented data roughly halved the accuracy drop (−31.4 → −16.0 points), clean accuracy rose to 83.47%, and on four typo types it never trained on it was 11 to 15 points more accurate than the baseline. Two confounds travel with that result: the augmented run took about 50% more optimizer steps, and no control duplicated rows without typos.
- Negative results. On the current run, temperature scaling made Choice calibration worse (ECE 0.0210 → 0.0371); isotonic recalibration brings it to 0.0113. Conformal sets met their 90% coverage target on all three heads, within finite-sample noise (Score: 89.85%). The constant-width Score interval averages about 2.6 stars once clipped to the valid range; adaptive intervals keep 90.11% coverage but cover 1-star reviews only 52.5% of the time, because 5-star reviews dominate the average. The speed and cost comparison against LLMs couldn’t run in the sandbox, so it is listed as not measured.
What I’d do next
- Re-run with per-example outputs saved, so the reliability diagrams and ECE under typos come from data rather than from mean confidence alone.
- Vary the perturbation seed, test on real user-typed text, add a pretrained encoder, and measure isotonic recalibration under typos.
- Run the speed and cost comparison against LLMs, and carry every new result into the write-up.
Links
- Code, findings and data licenses: github.com/mit37/BEMA (MIT)
- The paper: Can’t Hallucinate, Can Still Be Wrong
- The same instinct in a multi-model system: The Council
Verified numbers
| Metric | Value | Source |
|---|---|---|
| 77-way accuracy under two adjacent-character typos (BANKING77, n = 3,080) | 81.9% → 51.2% | mit37/BEMA/blob/d68f1a4/FINDINGS.md#L164-L173 |
| Mean confidence under the same typos | 0.791 → 0.612 | mit37/BEMA/blob/d68f1a4/FINDINGS.md#L164-L173 |
| Second domain under typos (CLINC150, 151 classes, n = 5,500) | 71.8% → 45.0% accuracy · 0.75 → 0.56 confidence | mit37/BEMA/blob/d68f1a4/CLINC150.md#L33-L49 |
| Binary head accuracy · ECE before → after temperature scaling | 97.85% · 0.0159 → 0.0046 | mit37/BEMA/blob/d68f1a4/results_log.csv#L35-L36 |
| 77-way head ECE · raw → temperature (worse) → isotonic | 0.0210 → 0.0371 → 0.0113 | mit37/BEMA/blob/d68f1a4/results_log.csv#L39-L41 |
| Accuracy drop under typos, before → after typo-augmented training | −31.4 → −16.0 pts | mit37/BEMA/blob/d68f1a4/FINDINGS.md#L460-L467 |
| Split-conformal coverage at a 90% target (binary · 77-way · score) | 90.91% · 98.83% · 89.85% | mit37/BEMA/blob/d68f1a4/CONFORMAL.md#L47-L53 |