Work / Research

BEMA

Stress-tested a “can’t-hallucinate” model design: 2 typos cut accuracy 82% → 51%, confidence only 0.79 → 0.61.

Status
research
Role
Independent research: specified the experiments, directed Claude Code agents through the build and write-up
Timeline
Sep 2026 →
  • Python
  • PyTorch
  • Hugging Face tokenizers (byte-level BPE, trained from scratch)
  • scikit-learn
  • NumPy

Live demo

Typo attack

The perturbation BEMA used, running live: two adjacent-character swaps. Then the measured cost of those two swaps, from the repo.

CleanHow do I locate my card?

What the model sawHowd o I locate my crad?

Two adjacent-character swaps per query, exactly as BEMA’s stress test applies them. mit37/BEMA/adversarial_stress_test.py#L193-L200 (opens in a new tab)

Why two swaps hurt so much

available →availablevaailable →vaailable

One swap splits one familiar token into four fragments the small tokenizer has barely seen. mit37/BEMA/FINDINGS.md#L394-L398 (opens in a new tab)

Two typos cut accuracy by 37.5%, but mean confidence by only 22.6% (relative, BANKING77).

Clean vs. two typos , full test splits.

BANKING77 · 77 classes · n = 3,080

Accuracy81.9% → 51.2%
Mean confidence79.1% → 61.2%

Under typos the model is 10.0 points more confident than it is accurate.

mit37/BEMA/FINDINGS.md#L164-L173 (opens in a new tab)

CLINC150 · 151 classes · n = 5,500

Accuracy71.8% → 45.0%
Mean confidence75.1% → 56.1%

Under typos the model is 11.1 points more confident than it is accurate.

mit37/BEMA/CLINC150.md#L46-L49 (opens in a new tab)

The fix, partly: train on typos

BANKING77 · same typo draws · n = 3,080

Accuracy81.9% → 50.5%
Mean confidence79.1% → 61.2%

Under typos the model is 10.7 points more confident than it is accurate.

mit37/BEMA/FINDINGS.md#L460-L467 (opens in a new tab)

No prediction is shown for your text: BEMA hasn’t exported per-example outputs yet, and this demo never guesses one. The typo is real; the numbers are the repo’s measured aggregates.

TL;DR

  • Typed-decision models answer in fixed-shape fields (a yes or no, a pick from a list, a number) in one forward pass, and the design is pitched as unable to hallucinate. I built an open, MIT-licensed reproduction of that interface to test what the claim leaves out.
  • On clean data it looks trustworthy. Under two ordinary typos, 77-way accuracy fell from 81.9% to 51.2% (n = 3,080) while mean confidence only fell from 0.79 to 0.61.
  • Calibration measured on clean text never sees that gap. The result reproduced on a second domain, and I report the negative results next to the positive ones.

The problem

A model with no free text to generate can’t invent a citation. It can still be wrong, and the question that matters is whether it knows when it is. The standard reliability metric, expected calibration error (ECE), is measured on clean test sets. Real inputs have typos. Fintech and quant teams live on that difference: a model that is confidently wrong under ordinary noise is worse than one that is uncertain.

What I built

I specified the work as a five-workstream research plan and directed Claude Code agents through the build, the experiments and the write-up. The model is trained from scratch on real human-labeled data only, never on LLM-generated labels.

  • Model. A small transformer encoder (3 layers, 6 heads, width 96) over a byte-level BPE tokenizer trained on the repo’s own training splits. Three typed heads read one shared state vector: Noul (binary: SMS spam), Choice (77-way banking intent: BANKING77) and Score (1–5 star ratings as regression). One forward pass answers all three.
  • Calibration. Temperature, Platt and isotonic recalibration, measured as top-label ECE with 10 bins. Multiclass isotonic recalibrates only the reported confidence and never changes the prediction.
  • Conformal prediction. Split conformal sets for all three heads, with the validation split halved so recalibration and conformal calibration never share data.
  • Stress tests. Two random adjacent-character swaps on every BANKING77 test query, spam-evasion tricks (spacing, leetspeak, case), cross-domain inputs, and a tokenizer-fragmentation audit.
  • Follow-ups. Typo-augmented retraining, tested on typo types it never saw; a standalone 151-class model on CLINC150, with out-of-scope rejection by confidence threshold; adaptive conformal intervals for Score; and an attention-pooling ablation across three training seeds.

Key decisions

  • Decision: real human labels only. Why: a test of hallucination can’t rest on labels a model generated. Trade-off: smaller datasets, and each one needs its own licensing check.
  • Decision: drop STS-B. Why: the agents flagged mixed licensing terms across its sub-sources, and I made the call to cut it; the repo keeps a per-dataset licensing record. Trade-off: one fewer regression benchmark.
  • Decision: report the pretrained-encoder comparison as not run rather than approximate it. Why: the build environment blocked the model hubs, so every encoder is trained from scratch, and the plan’s standing rule was never to fake an experiment that couldn’t run. Trade-off: a small model, and the biggest open question stays open.
  • Decision: report what didn’t work. Why: a calibration study that hides its failures isn’t one. Trade-off: the write-up reads less like a highlight reel, which is the point.

The hard part

Is the typo collapse a property of the architecture, or an artifact of my tokenizer? The 81.9% → 51.2% number is striking, and it would be easy to over-claim it.

The audit showed that byte-level BPE never produces an unknown token, but a two-character swap often splits one clean token into several weak ones: “available” becomes “vaailable”, which tokenizes as v, aa, ila, ble. Four of five sampled queries fragmented by two to five extra tokens. So I split the finding into two claims:

  1. Confidence doesn’t track accuracy under ordinary noise. A roughly 23% relative confidence drop against a roughly 37% relative accuracy drop. This is the headline, and it should hold whatever the tokenizer.
  2. The size of the collapse is explicitly not claimed for pretrained tokenizers with larger vocabularies. The experiment that would settle it, a pretrained encoder, couldn’t run in the build sandbox, so it stays listed as unresolved rather than approximated.

A second catch came from an agent-run review across correctness, statistics, licensing, reproducibility and security, with adversarial verification of each finding. It found that the multiclass isotonic path overwrote a probability in place and re-took the argmax, which could silently flip predictions. The fix derives predictions from the raw softmax and recalibrates only the reported confidence.

Results

  • Clean data. Noul reaches 97.85% accuracy, and temperature scaling cuts its ECE from 0.0159 to 0.0046. Choice scores 81.88% on BANKING77’s full 3,080-query test set. Score reaches a test MAE of 0.187 with Pearson r = 0.53, the lowest of three training seeds (0.53 to 0.58).
  • Under typos. Choice accuracy falls from 81.9% to 51.2% while mean confidence only falls from 0.79 to 0.61. On CLINC150, a separate 151-class model falls from 71.8% to 45.0% while confidence falls only from 0.75 to 0.56 (n = 5,500).
  • Mitigation. Retraining on typo-augmented data roughly halved the accuracy drop (−31.4 → −16.0 points), clean accuracy rose to 83.47%, and on four typo types it never trained on it was 11 to 15 points more accurate than the baseline. Two confounds travel with that result: the augmented run took about 50% more optimizer steps, and no control duplicated rows without typos.
  • Negative results. On the current run, temperature scaling made Choice calibration worse (ECE 0.0210 → 0.0371); isotonic recalibration brings it to 0.0113. Conformal sets met their 90% coverage target on all three heads, within finite-sample noise (Score: 89.85%). The constant-width Score interval averages about 2.6 stars once clipped to the valid range; adaptive intervals keep 90.11% coverage but cover 1-star reviews only 52.5% of the time, because 5-star reviews dominate the average. The speed and cost comparison against LLMs couldn’t run in the sandbox, so it is listed as not measured.

What I’d do next

  • Re-run with per-example outputs saved, so the reliability diagrams and ECE under typos come from data rather than from mean confidence alone.
  • Vary the perturbation seed, test on real user-typed text, add a pretrained encoder, and measure isotonic recalibration under typos.
  • Run the speed and cost comparison against LLMs, and carry every new result into the write-up.

Verified numbers

MetricValueSource
77-way accuracy under two adjacent-character typos (BANKING77, n = 3,080)81.9% → 51.2%mit37/BEMA/blob/d68f1a4/FINDINGS.md#L164-L173
Mean confidence under the same typos0.791 → 0.612mit37/BEMA/blob/d68f1a4/FINDINGS.md#L164-L173
Second domain under typos (CLINC150, 151 classes, n = 5,500)71.8% → 45.0% accuracy · 0.75 → 0.56 confidencemit37/BEMA/blob/d68f1a4/CLINC150.md#L33-L49
Binary head accuracy · ECE before → after temperature scaling97.85% · 0.0159 → 0.0046mit37/BEMA/blob/d68f1a4/results_log.csv#L35-L36
77-way head ECE · raw → temperature (worse) → isotonic0.0210 → 0.0371 → 0.0113mit37/BEMA/blob/d68f1a4/results_log.csv#L39-L41
Accuracy drop under typos, before → after typo-augmented training−31.4 → −16.0 ptsmit37/BEMA/blob/d68f1a4/FINDINGS.md#L460-L467
Split-conformal coverage at a 90% target (binary · 77-way · score)90.91% · 98.83% · 89.85%mit37/BEMA/blob/d68f1a4/CONFORMAL.md#L47-L53
m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h