Research / Essay

Uncorrelated Failure Modes

Abstract

The design rule behind most of what I build: never trust one source of truth. Independent methods vote, deterministic code does the arithmetic and keeps the books, every fact carries its provenance, and a person decides whenever the checks disagree. This essay explains the rule through my own systems, including where it has already failed me: checks that were supposed to be independent, and weren’t.

A model I built to classify banking questions was right 82% of the time on clean test data, and its confidence said so. Then I added two typos to every question, the kind anyone makes typing fast. It was now right about half the time, 51%, and its average confidence had only slipped from 0.79 to 0.61. Nothing inside the model announced the change. The one signal it offered about its own reliability had quietly stopped tracking reality.

That result is the cleanest statement I have of a problem I keep designing around. A single check fails silently. It fails the same way every time on the same kind of input, and you cannot see its blind spots from inside it, because they are its blind spots. So I don’t let one method be the only thing standing between a model and a consequence. I want several checks whose failures don’t line up, and a person on hand for the moments when they disagree.

I call the goal uncorrelated failure modes. It sounds like a statistics term because it is one. It is also the most practical rule I know for building software that touches money, courts or people. (On the systems below, I architected and specified the design and directed AI coding agents to write the code. The design rules are the part this essay is about.)

Independence is the point, not the count

The instinct is to add checks. Three checks sound better than one. But three checks that fail on the same inputs are one check run three times.

In AccountWard, the accounting engine I designed for court-supervised fiduciaries, three separate methods vote on how to categorize every bank transaction. The first is a set of deterministic payee rules. The second is a nearest match against examples a human has already approved. The third is a model vote that has to pass a strict JSON gate or abstain. They are different in kind on purpose: a rule fails when the payee is new, a similarity match fails when the description is unusual, and a model fails in ways that are harder to predict. When all three agree, the item gets 0.98 confidence. When they split two to one, it gets 0.75.

The history of this idea is humbling. In a 1986 experiment, Knight and Leveson had 27 versions of the same program written independently from one specification at two universities, then ran all of them against a million tests. The versions failed together far more often than independence would predict (Knight & Leveson, 1986). Different people, same hard cases. Language models repeat the lesson: a 2025 study of more than 350 models found that when two models both got a question wrong, they often gave the same wrong answer, and that larger, more accurate models had highly correlated errors even across providers (Kim et al., 2025).

So “use a different model” is not a design. The question I ask is what each check actually looks at, and whether that differs from what the others look at. A second model from a second lab that reads the same text with the same prompt is a weaker second opinion than a string comparison that reads nothing but the source document.

Humans get the split vote

If the checks agree, automate. If they disagree, escalate. That is the whole routing rule, and it has one important consequence: a vote never decides the answer on its own.

In AccountWard, anything under 0.9 confidence goes to a human review queue, least certain first, so every two-to-one split reaches a person. In Warrant, the estate-settlement engine I built with Fluxxion88 at a hackathon, disagreements between two independent form-mapping runs go to a human worklist, never to a majority. A mentor at that hackathon gave the example that made the rule stick. Average a $740,000 appraisal and a $760,000 appraisal and you land exactly on a $750,000 statutory threshold: the one answer that is guaranteed to be wrong in the most consequential way. A vote averages the disagreement away. A person resolves it.

The same rule shaped Concord, my due-diligence engine, where every analyst’s findings are challenged by a model from a different provider whenever a second provider is connected, and the risk committee records each dispute with both positions instead of a winner. The dissent is the product. It tells the reader exactly where to look.

Rules check what rules can check

Most of the checking in my systems isn’t done by models at all. Some questions have exact answers, and exact answers belong to code.

In NEXUS, a hackathon engine for wargame adjudication and supply-chain resilience, five language-model agents argue about what a disruption means. None of them touches the arithmetic. A deterministic cascade model runs 10,000 Monte Carlo trials between their turns and reports a failure probability with a confidence interval. In AccountWard, the money math lives in a pure TypeScript engine that works only in integer cents, with no floats and no I/O, so the entire accounting can be recomputed from the transactions and checked to the penny. BlueCollarPal, a procurement tool for trade contractors, was built on one sentence: models extract; deterministic code controls money.

The strongest version of this is to keep the model out of the runtime entirely. My teammate Fluxxion88’s half of Warrant + Forge, a form compiler called Forge, uses a model once per form at build time to propose how fields bind to records. A human approves the binding, and after that every fill is deterministic: one IRS form for five sample estates, filled in 40 to 42 milliseconds each, with zero model calls. That was our thesis for the demo: we don’t fill forms with AI, we use AI to build form-fillers.

Determinism buys something models can’t: the same answer twice. In a small test of Concord’s citation checker, I asked an LLM judge the same question five times per citation. On 2 of 10 citations its verdict changed between asks, and it called one fabricated paraphrase present in 3 of 5 tries. The deterministic tie-out returns the same grade every time, which means a reviewer six months later can replay the file and get the same register. Concord recomputes a 102-entry risk register in about 131 milliseconds with no model calls.

Provenance on every fact

A check can only be audited if you can see what it checked against. So every fact in my systems carries where it came from.

Warrant’s rule is in its name: a model may propose a fact, but it cannot assert one. Every fact needs a warrant, which is a verbatim quote that a deterministic matcher finds in a document on file, a recorded derivation, or a record path. A fact without one is quarantined, visible to a person and invisible to every rule. AccountWard’s legal rulebook holds 106 records, forms, statutes and county rules, and each one cites its primary source and the date it was retrieved. Its audit log is append-only and hash-chained, so a history cannot be rewritten without the chain showing it. Even BEMA, the research project behind the typo result, keeps a record of how each dataset’s license was verified, and it dropped a dataset when that record came back unclear.

Provenance doesn’t prevent errors. It makes them findable. When something is wrong, I want to trace it to the source that said so in one step.

Where it breaks

This rule has already failed me, and the ways it failed are the most useful part of it.

Shared inputs. In Concord’s citation benchmark, a string matcher and an LLM judge are about as different as two checkers get. They agreed on 19 of 20 planted citations, and were both wrong on 2 of those 19: faithful quotes from a scanned document whose OCR had mangled the text. The checkers were independent. The input was shared, and the input was what failed. The fix was a change to the tie-out, not another model.

Shared design. In AccountWard’s default build, the payee rules and the stand-in for the model vote both read the payee text. Three methods, partly one signal. It is written into the case study as a caveat, and it stays there until accuracy by vote pattern is measured.

Shared family. Warrant’s two mapping runs came from two models by the same lab. They agreed on 87 of 94 fields, and all 3 known errors sat inside the 7 disagreements. That is encouraging, and it is not proof: nobody audited the 87 agreements, and one error was found by looking at a disagreement.

The boundary of a check. Warrant’s verifier caught 24 of 24 number substitutions in quoted text. It caught 0 of 41 cases where a genuine quote was attached to a wrong value, because it compares the quote to the document and never the value to the quote. That zero went into the published benchmark. A check you don’t know the edges of is worse than no check, because you trust it past its edges.

The rule, restated

Uncorrelated is not a label I get to put on a system. It is a property I have to measure, pair by pair: what does each check look at, and how often does it fail on the same items as the others? Until I have measured it, agreement is something to log and sample, and disagreement is the signal that sends an item to a person.

In practice it comes down to four habits. Build checks that differ in kind, not just in vendor. Let code do the arithmetic and keep the books. Attach a source to every fact. Give every split to a human. None of this makes a system right. It makes a system wrong in ways I can see.

m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h