Research / Paperposition paper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systems
Abstract
A language model is most dangerous when it is confidently wrong, and its own confidence is a weak guard. This position paper describes an architecture I built nine times between April and August 2026. Agents do the work, deterministic checks that no model can vote on test it, a council of models from other labs challenges it, and a human takes every split decision. From those builds come three working theories: adjudication beats synthesis, the model that writes should not check its own work, and agreement is not proof. Only the third has data behind it so far. In a two-model cross-check of 94 form-field mappings, all 3 known errors sat inside the 7 disagreements, though the agreements were never audited. On 20 planted citations, a string matcher and an LLM judge agreed on 19 and were both wrong on 2 of those, so their failures overlapped instead of cancelling. Recent work finds that errors correlate across model providers too. I therefore treat independence as a property to measure, not a feature to buy, and propose an evaluation on three tasks with checkable ground truth: error rates when checkers agree and when they disagree, adjudication against synthesis, and self-review against cross-lab review.
Key findings
- In a two-model cross-check of 94 government-form field mappings, all 3 known errors sat inside the 7 disagreements. The 87 agreements were never audited, so this is a floor, not a rate.
mit37/warrant/docs/BENCHMARKS.md#L100-L112 - A deterministic citation checker and an LLM judge agreed on 19 of 20 planted citations and were both wrong on 2 of those 19: their errors overlapped instead of cancelling.
local: concord/bench/uncorrelated-result.json - Asked the identical question five times, the same LLM judge gave more than one verdict on 2 of 10 citations, and called a fabricated paraphrase present in 3 of 5 tries.
local: concord/bench/stability-result.json - Of the three working theories, only “agreement is not proof” has data behind it. For the other two this paper proposes the experiment and claims no result.
Introduction: why one model’s confidence is not enough
Language models are trained and scored in ways that reward a confident guess over an admission of uncertainty (Kalai et al., 2025). The result is familiar: fluent answers that are wrong, delivered in the same voice as the ones that are right. The natural guard is the model’s own confidence, and it is weak. In a companion study, even a model built to answer only in fixed-schema fields, with textbook calibration on clean data, stopped tracking its own accuracy once each input carried two ordinary typos: accuracy fell from 81.9% to 51.2% while confidence fell only from 0.791 to 0.612 (Mittal, 2026).
Asking a model to grade its own work does not fix this. LLM judges show position, verbosity and self-enhancement biases (Zheng et al., 2023). Models can recognize their own outputs and prefer them (Panickssery et al., 2024). Without external feedback, a model’s attempt to correct its own reasoning can make it worse (Huang et al., 2023).
The obvious next step is to ask several models and merge their answers. That has its own failure mode. If the models share a blind spot, merging averages the error in, and agreement among them looks like confirmation. I saw a version of this in the first system in this line. NEXUS, a hackathon wargame-adjudication engine, had five agents debate the output of a Monte Carlo cascade model. The first version’s parameters were degenerate, so nearly every trial failed (0.958 in the first smoke test), and the agents argued fluently over a number that carried no information.
This paper makes a design argument and defines the experiment that would test it. It describes an architecture I built, in nine versions, between April and August 2026. I specified each build and directed AI coding agents (Claude Code, Codex and Cursor) to write the code. It states three working theories that came out of those builds. It reports the two small measurements that exist, with their caveats. For the rest, it proposes an evaluation and claims no result that has not been run.
Related work
Ensembles and correlated failure. Disagreement among independently trained models is an established uncertainty signal: deep ensembles use it to flag inputs a single network would answer overconfidently (Lakshminarayanan et al., 2016). Its weak point is the word “independently”. In a classic software experiment, 27 versions of a program written independently to the same specification failed together on far more tests than independence would predict (Knight & Leveson, 1986). Language models show the same pattern. In a study of more than 350 models, when two models both got a question wrong on one leaderboard dataset, they gave the same wrong answer 60% of the time, and larger, more accurate models had highly correlated errors even across distinct architectures and providers (Kim et al., 2025). Any design that counts agreement as evidence has to budget for this.
Consistency within one model. Self-consistency samples several reasoning paths and takes the majority answer (Wang et al., 2022). SelfCheckGPT flags a claim as a likely hallucination when stochastically sampled answers diverge (Manakul et al., 2023). Both use disagreement as the signal, as this paper does, but among samples from a single model. The architecture here asks for disagreement across models, across labs and across methods, including methods with no model in them.
Debate. Debate was proposed as a safety technique in which two agents argue and a human judges (Irving et al., 2018). Multi-agent debate among language models improves factuality and reasoning (Du et al., 2023), and debate can counter a single model’s tendency to stop generating new thoughts once it is confident in its answer (Liang et al., 2023). Debate protocols usually drive toward one consensus answer. The protocol here keeps the losing position on the record and sends unresolved splits to a person.
Aggregation. Mixture-of-Agents layers agents so that each one refines the outputs of the layer before, and it reported leading results on open chat benchmarks (Wang et al., 2024). It is the strongest version of the synthesis design that this paper’s first theory argues against, for tasks with checkable answers. Its results are the main reason that theory is stated as a hypothesis.
LLM judges and panels. Strong judges agree with human preferences over 80% of the time, about as often as humans agree with each other (Zheng et al., 2023). A panel of smaller judges drawn from different model families outperformed a single large judge across six datasets, with less intra-model bias, at over seven times lower cost (Verga et al., 2024). That is the closest prior result to this paper’s second theory.
Deferral and calibration. Learning-to-defer methods train a model to decide when to hand a decision to a human expert (Mozannar & Sontag, 2020). This protocol uses a blunter rule: every disagreement is deferred. Temperature scaling and related methods make a single model’s confidence better calibrated on held-out data (Guo et al., 2017), and large models are well calibrated on multiple-choice and true/false questions posed in the right format (Kadavath et al., 2022). The companion study shows how quickly that can fail once inputs get noisy.
Architecture
The design separates doing the work from checking it, and checking by models from checking by rules.
Agreement is logged, not trusted: a sample of agreed items is audited, because checkers can fail together. Items where agreement is weak join the human queue too, least confident first.
Agent layer
Agents do the work: an analyst desk reads a data room, a mapper binds a form field to a record, a classifier assigns a transaction to a schedule. The rule that makes everything downstream possible is that every claim carries its evidence. In the diligence engine, each finding must cite a verbatim quote. In the estate engine, a model may propose a fact but not assert it: the fact needs a warrant, meaning a quote a matcher finds in a held document, a recorded derivation, or a record path. Anything without a warrant is quarantined, visible to a person and invisible to every rule.
Deterministic checks
Before any model judges anything, rules do what rules can do. A citation tie-out decides whether a quote appears in the document it cites: normalized containment first, then n-gram coverage (4-grams for quotes of eight or more tokens, bigrams below that), with coverage of at least 0.85 counting as verified. In the estate engine, a digit guard also requires every number in a quote to appear in the document. It was added after a test showed that changing one figure in an eleven-token quote still left coverage at 0.875, above the bar. Invariants play the same role for money: in the accounting engine, charges must equal credits in every period, and every transaction must land on a schedule. No model votes on any of this. A string is in the document or it is not.
Council layer
A challenger then reviews each piece of work. In the reference build, Concord, each analyst desk faces a challenger model that returns one of three verdicts: upheld, challenged or missed. The analyst answers on its own model. The challenger is selected to avoid the analyst’s provider: whenever a second provider is connected, the analyst’s own provider is excluded outright, and six tests check that across every provider pairing in the model catalog. With only one provider connected, the seat falls back to the same vendor. A devil’s advocate that avoids the analysts’ providers follows.
Adjudication protocol
The last stage routes each item by agreement.
- Agreement is logged, not trusted. Both verdicts stay on the record, and a sample of agreed items is audited, because checkers can fail together (see “Agreement is not proof” below).
- Disagreement goes to a person, with both positions. Concord’s risk committee records each dispute with both sides, and its investment-committee memo must carry a dissent section.
- Votes never set a value. When two verified sources disagree on a number, that is a conflict for a human to resolve. A mentor at the estate-settlement hackathon gave the example that made the rule concrete: average a $740,000 appraisal and a $760,000 appraisal and you land exactly on a $750,000 statutory threshold, the worst possible answer.
- Confidence comes from the agreement pattern. In the accounting engine, three of three methods agreeing gives 0.98, two plus an abstention 0.92, a two-to-one split 0.75 and a single vote 0.6. Anything under 0.9 goes to the review queue, least certain first, so every split reaches a person.
- A failed seat is not a vote. When a provider fails mid-run, its placeholder output must never stand in for analysis. Concord learned this the hard way (see Applications).
Nine builds
The architecture was not designed in one pass. It is where nine builds converged, and the direction of travel was consistent: from synthesis toward adjudication, and from model-only checking toward deterministic checks in front of the models.
| # | Build | When | Protocol | Model diversity |
|---|---|---|---|---|
| 1 | NEXUS, wargame adjudication | Apr 2026 | Sequential agent debate, then an arbiter’s synthesis, with a Monte Carlo check | one vendor |
| 2 | Cross-vendor M&A “grand council” | May 2026 | Per-vendor agent swarms, then cross-vendor arbitration and a red-team loop | four vendors in the live run |
| 3 | Coding-agent review council | Jun 2026 | Five reviewers, then a judge’s synthesis | mock provider |
| 4 | BlueCollarPal takeoff | Jul 2026 | Estimator against skeptic; a deterministic reconciler sends challenged lines to a human | one model |
| 5 | Redl Council pane | Jul 2026 | Layered synthesis pyramid, 4 → 3 → 1 | one local model by default; per-seat routing |
| 6 | Quorum | Jul 2026 | Redl’s pyramid with diligence roles | one model |
| 7 | Concord | Jul–Aug 2026 | Different-vendor challenger, deterministic tie-out, evidence auditor, mandatory dissent | cross-vendor at the challenge stages |
| 8 | Warrant cross-check | Jul 2026 | Two independent mapping runs; human adjudications stored as data | one vendor, two models |
| 9 | AccountWard classifier | Jul–Aug 2026 | Three-method vote; below 0.9 goes to a human | three methods |
Three working theories
Adjudication beats synthesis
When several models answer, asking one of them to merge the answers averages the errors in. A merged answer inherits every participant’s mistakes in diluted form and hides which participant made them. Asking a judge to pick one answer, justify it against the others and keep the dissent on the record preserves the information that something is contested.
Status: a hypothesis. Both styles exist in my builds. Redl’s Council is a synthesis pyramid and Concord adjudicates. The two have never been compared on the same task. Mixture-of-Agents’ benchmark results are good evidence that synthesis works well for open-ended generation. The theory is narrower: for decisions with checkable answers, a correct minority answer should survive more often under adjudication.
The model that writes shouldn’t check its own work
Self-review shares the author’s blind spots. The published evidence on self-preference and on self-correction without external feedback points the same way. So the checker should differ from the author in whatever way is cheapest to guarantee: a different lab, a different model, or no model at all.
Status: enforced in code, not measured. Concord excludes the analyst’s provider from the challenger seat whenever it can. There is also a counter-anecdote. In BlueCollarPal’s blueprint takeoff, a skeptic running on the same model as the estimator still caught the estimator’s arithmetic error on a planted panel schedule. That is one case, and it shows that a same-model reviewer with a different role and context is not useless. And because errors correlate across providers as well (Kim et al., 2025), “a different lab” is a heuristic for independence, not a guarantee of it.
Agreement is not proof
Models share training data, and checkers share inputs. When two checkers are each wrong with probability and , and their errors are correlated with coefficient , the chance that both are wrong on the same item is
For two checkers that are each wrong 10% of the time, independence gives 1 item in 100 where both fail. A correlation of 0.5 raises that to 5.5 in 100. Agreement among correlated checkers is weak evidence, and the weakness is invisible unless someone audits the agreements.
Status: supported by the only data that exists, which follows.
Proposed evaluation
Preliminary evidence
Two small benchmarks were built for other reasons: one to calibrate a form-mapping pipeline, the other to test a citation checker. Neither is the experiment below. They are reported because they are the whole evidence base so far, and each comes with its caveats.
| Two runs | Fields | Known errors | Known error rate |
|---|---|---|---|
| Runs agree | 87 | 0 | unaudited |
| Runs disagree | 7 | 3 | ≥ 3 of 7 (42.86%) |
Warrant (estate-settlement hackathon build, July 2026). Two runs from the same model family mapped fillable fields on four government forms. Known errors are the human adjudications, rewound before comparing.
- One error was found by inspecting a disagreement, so the setup favors this result.
- Same prompt, same evidence, same model family: a sibling, not an independent checker.
- Ground truth is partial; every error rate is a lower bound.
In the Warrant cross-check, disagreement was a useful filter: 3 of the 7 disagreements were known errors, and none of the known errors sat among the agreements. But the 87 agreements were never audited, one of the three errors was found by inspecting a disagreement, and the two runs came from the same model family. The benchmark’s own write-up puts it plainly: a sibling model is a second opinion, not an independent check.
| Outcome | Citations | Cases |
|---|---|---|
| Both right | 17 | — |
| Both wrong (agreed, and wrong) | 2 | O1, O2 (OCR-degraded, faithful) |
| Disagree: tie-out right, LLM wrong | 1 | F4 (faithful, whitespace variant) |
| Disagree: LLM right, tie-out wrong | 0 | — |
Concord (M&A diligence engine, August 2026). A deterministic tie-out and one LLM judge (gpt-5.6-luna) each decided whether a quoted passage appears in its cited synthetic document. “Right” means the verdict matched the planted truth.
- Small n on synthetic documents. The two shared failures came from the input (OCR-damaged text), not from either checker, which is exactly the kind of correlation diversity of checkers cannot remove.
- Run before the OCR fold was added to the tie-out; a re-run may change the numbers.
- The result files are not yet committed to the Concord repository.
The Concord comparison pitted two checkers as different as checkers get, a string matcher and a language model, against 20 planted citations. They agreed on 19, and 2 of those agreements were wrong. Both were faithful quotes from an OCR-damaged document, which both checkers rejected. The shared failure came from the input, not from either checker, and that is the kind of correlation that diversity among checkers cannot remove. Neither checker accepted a fabrication. The fix for the OCR case was a change to the tie-out (a fold for degraded text, capped at a lower grade), not another model.
| Measure | Result |
|---|---|
| Citations with more than one verdict | 2 of 10 |
| Fabricated paraphrase (P1) judged PRESENT | 3 of 5 trials |
| Faithful quote (F4) judged ABSENT | 4 of 5 trials |
Identical prompt, five trials per citation, ten citations. A deterministic tie-out returns the same verdict on every repeat by construction.
- One model, ten citations, five trials: enough to show instability exists, not to estimate its rate.
Asked the same question five times, the LLM judge changed its verdict on 2 of 10 citations. A judge that disagrees with itself is one more reason to measure a checker before counting its vote.
Tasks
The evaluation needs ground truth that can be checked without a model. Three tasks from the builds above have it:
- Citation verification. Concord’s planted corpus: 23 quotations against 4 synthetic diligence documents, 8 faithful but damaged (line breaks, smart quotes, whitespace, truncation, OCR noise) and 15 fabricated (absent, wrong document, altered digits, flipped negations, swapped entities, paraphrases and attacks on the OCR fold). Extend it with fresh cases written by someone who has not seen the checkers.
- Form-field mapping. Warrant’s 94 fields mapped by two runs, re-run with models from two or three labs, and with every agreement audited by hand so that errors among agreements are counted, not assumed absent.
- Transaction classification. AccountWard’s synthetic corpus of 684 transactions over 24 months, with 12 planted hard cases, scored by vote pattern: three of three, two plus an abstention, two to one, and a single vote.
Conditions
For each task, an author model produces the work, and it is checked by four kinds of reviewer: the author itself, a sibling model from the same lab, a model from a different lab, and a deterministic check where one exists. Two or three labs are enough to start.
Measures
The primary measure is, for each pair of checkers, the error rate when they agree against the error rate when they disagree, and the ratio of the two:
A large lift means disagreement is a good trigger for human review. A small means agreement can be automated. The other measures:
- Error correlation. The φ coefficient between each pair of checkers’ error indicators, which plugs directly into the formula above.
- Correct-minority survival (for adjudication against synthesis). On items where exactly one council member is right, how often does the final answer end up right under each aggregation protocol?
- Reviewer recall and false-alarm rate (for self-review against cross-lab review), against the planted and audited errors.
- Verdict stability. Every model verdict repeated at least five times with identical prompts.
- Cost, latency and failure rate per decision, including provider outages.
Protocol
Fix the analysis plan before running. Blind reviewers to which model wrote the work. Randomize the order of candidate answers shown to any judge, given the known position bias. Report Wilson intervals, because the counts will be small. Publish every run, including failed ones, with the result files committed next to the code.
What would change my mind
- If is close to on these tasks, disagreement is not a useful review trigger, and the architecture’s routing rule fails.
- If a different-lab reviewer’s errors correlate with the author’s about as much as a sibling’s do, the second theory reduces to “use a different model”, and the lab boundary buys nothing.
- If synthesis preserves correct minority answers as often as adjudication, the first theory is wrong for these tasks.
Applications
AccountWard (case study) applies the protocol to money. Three separate methods vote on each transaction’s category: payee rules, a nearest match against human-approved examples, and a model vote that must pass a strict JSON gate or abstain. The agreement pattern sets the confidence, and anything under 0.9 goes to a person. One caveat matters for this paper: in the default build, the rules and the stand-in for the model vote both read payee text, so the votes are only partly independent, and accuracy by vote pattern has not been measured. It is task 3 above.
Concord (case study) is the reference build: seven stages, eleven parallel analyst desks, a different-vendor challenger for each and a deterministic tie-out on every quote. In one complete live run over a synthetic five-document data room, all 11 challenges ran on a different vendor’s model from the analyst they checked, and the run returned 152 cited findings, 100 of which tied out cleanly. Diversity has an operating cost. In a later live re-run, one vendor’s credit ran out during cross-examination, and every rebuttal came back as a failure placeholder. The placeholder was a non-empty string, so it overwrote eleven finished analyses and flowed into the committee as if it were analysis. The register came back empty. The fix normalizes the failure marker, checks for it wherever output is carried forward, stops calling a provider once it is down, and pins the behavior with an outage test. More vendors means more ways for a seat to fail, and a failed seat must never count as a vote.
Warrant (case study), my half of Warrant + Forge, built with Fluxxion88 at an estate-settlement hackathon, takes the strictest line. Facts need warrants, rules are evaluated in three-valued logic so a missing fact blocks a decision instead of defaulting, and disagreements between two mapping runs go to a human worklist, never to a vote. The build plan included a model council, listed it second to cut if time ran short, and set a condition: with only one provider, cross-model challenge would be theatre, and it was better to cut it than fake it. It was cut.
Redl (case study) runs the other design. Its Council pane is a layered synthesis pyramid: four specialists, then a proposer, a critic and a red team, then a judge, each layer building on the one below. Every seat defaults to the same local model and can be routed to another provider. It is the synthesis arm that the first theory would test against. The July 2026 build has no side-by-side view that runs one prompt across several models and highlights where they diverge; that view is planned.
Limitations
- Small evidence. Two benchmarks with 94 fields and 20 citations, synthetic documents, same-family models in one case, and unaudited agreements. Nothing here estimates a rate.
- Diversity is not independence. Shared training data, shared prompts and shared inputs all correlate errors. The OCR case above is a shared-input failure no council can vote away.
- Cost and latency. Every extra seat is another call. The one complete live run above took about 50 minutes and cost $4.995 for a five-document synthetic data room.
- Operational fragility. More providers means more outages, rate limits and credit failures, each of which must be handled so that a failure is never mistaken for a verdict.
- The human queue is the cost center. Routing every split to a person is only affordable when splits are rare. The evaluation’s measure of decides how much can be automated.
- Provenance of the evidence. The code in every build was written by AI coding agents under my direction, and some Concord benchmark files are not yet committed.
Conclusion
The claim of this paper is modest and, I think, durable. A second model is not a second opinion unless its errors are independent of the first model’s, and independence has to be measured. Until it is, the safe design treats agreement as something to log and sample, and disagreement as the signal that routes an item to a person. The architecture that follows is simple to state. Agents do the work with their evidence attached. Rules check what rules can check. Checkers from other labs challenge the rest, and a human takes every split. The first two theories behind it are still hypotheses, and the evaluation above says how each could fail. The third, that agreement is not proof, is already visible in the smallest data I have.
References
- Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I. (2023). Improving Factuality and Reasoning in Language Models through Multiagent Debate. arXiv:2305.14325
- Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017). On Calibration of Modern Neural Networks. arXiv:1706.04599
- Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798
- Irving, G., Christiano, P., and Amodei, D. (2018). AI safety via debate. arXiv:1805.00899
- Kadavath, S., et al. (2022). Language Models (Mostly) Know What They Know. arXiv:2207.05221
- Kalai, A. T., Nachum, O., Vempala, S. S., and Zhang, E. (2025). Why Language Models Hallucinate. arXiv:2509.04664
- Kim, E., Garg, A., Peng, K., and Garg, N. (2025). Correlated Errors in Large Language Models. arXiv:2506.07962
- Knight, J. C., and Leveson, N. G. (1986). An experimental evaluation of the assumption of independence in multiversion programming. IEEE Transactions on Software Engineering 12(1), 96–109. doi:10.1109/TSE.1986.6312924
- Lakshminarayanan, B., Pritzel, A., and Blundell, C. (2016). Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. arXiv:1612.01474
- Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Shi, S., and Tu, Z. (2023). Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate. arXiv:2305.19118
- Manakul, P., Liusie, A., and Gales, M. J. F. (2023). SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. arXiv:2303.08896
- Mittal, M. (2026). Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noise. Working draft
- Mozannar, H., and Sontag, D. (2020). Consistent Estimators for Learning to Defer to an Expert. arXiv:2006.01862
- Panickssery, A., Bowman, S. R., and Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv:2404.13076
- Verga, P., Hofstätter, S., Althammer, S., Su, Y., Piktus, A., Arkhangorodsky, A., Xu, M., White, N., and Lewis, P. (2024). Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. arXiv:2404.18796
- Wang, J., Wang, J., Athiwaratkun, B., Zhang, C., and Zou, J. (2024). Mixture-of-Agents Enhances Large Language Model Capabilities. arXiv:2406.04692
- Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. (2022). Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171
- Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685