Work / AI Systems
The Council
Agents do the work, other models check it. Built nine ways, then measured, with the negative results reported.
- TypeScript
- Rust
- Python
- Anthropic API
- OpenAI-compatible APIs
- Vitest
Live demo
Council playground
Several answers to one question. The disagreement layer finds the claims they split on, and the threshold decides what a human has to look at.
QuestionA $1,000 deposit earns 5% simple interest a year. What is the balance after 3 years?
A
Simple interest pays 5% of the original $1,000 each year: $50 × 3 = $150. The balance is $1,150.00.
B
Interest accrues on the principal only, with no compounding, so the deposit earns $50 a year. After three years it stands at $1,150.
Cdissents on 2
Compounding annually: $1,000 × 1.05³ = $1,157.63.
D
Three years of $50 in interest is $150, for a balance of $1,150.00.
Worst ∆0.33
2 of 2 claims go to a human.
| Claim | Answers, grouped | Agreement | Decision |
|---|---|---|---|
| Balance after 3 years | 1,150 A B D1,157.63 C | 75% ∆ 0.25 | human majority is right |
| Interest method | simple A Bcompound C | 67% ∆ 0.33 | human majority is right |
One answer compounds when the question says simple interest. The other three agree, and the split still reaches a human at the default threshold.
Scripted examples written for this demo. They are not output from any AI model, and the labels A to E are placeholders.
TL;DR
- Between April and August 2026 I built the same idea nine times: agents that do the work, and a council of other models that checks it.
- The design moved from synthesis (merge everyone’s answer) to adjudication (keep dissent on the record, challenge each analyst with another vendor’s model, check every citation deterministically, send split votes to a human).
- Then I measured the premise. Where the data cut against it, I report that too: agreement is evidence, not proof.
The problem
A single model is most dangerous when it is confidently wrong, and asking it to check its own work shares its blind spots. The common fix, asking several models and merging their answers, averages the errors in instead of surfacing them. I wanted an architecture where disagreement is treated as a signal to act on.
What I built
I specified each build and directed AI coding agents (Claude Code, Codex and Cursor) to write the code. The lineage:
| # | Build | When | Protocol |
|---|---|---|---|
| 1 | NEXUS, wargame adjudication | Apr 2026 | Sequential agent debate, then an arbiter’s synthesis, with a Monte Carlo check |
| 2 | Cross-vendor M&A “grand council” | May 2026 | Per-vendor agent swarms, then cross-vendor arbitration and a red-team loop |
| 3 | Coding-agent review council | Jun 2026 | Five reviewers, then a judge’s synthesis |
| 4 | BlueCollarPal takeoff | Jul 2026 | Estimator against skeptic; a deterministic reconciler sends challenged lines to a human |
| 5 | Redl Council pane | Jul 2026 | Layered synthesis pyramid, 4 → 3 → 1 |
| 6 | Quorum | Jul 2026 | Redl’s pyramid with diligence roles |
| 7 | Concord | Jul–Aug 2026 | Cross-examination by a different-vendor challenger, deterministic tie-out, evidence auditor, mandatory dissent |
| 8 | Warrant cross-check | Jul 2026 | Two independent mapping runs, with human adjudication stored as data |
| 9 | AccountWard classifier | Jul–Aug 2026 | Three-method vote; below 0.9 confidence goes to a human |
The reference build is Concord’s seven-stage diligence pipeline. Parallel analyst desks run first. A cross-examination round seats each challenger on a different provider from its analyst whenever a second provider is connected. Then come a devil’s advocate, a deterministic citation tie-out with no model in the loop, an evidence auditor, a risk committee that records disputes with both positions, and an investment-committee memo with a mandatory dissent section.
Key decisions
- Decision: adjudicate, don’t synthesize. Why: merging answers averages errors in; picking and justifying keeps the dissent visible. Trade-off: more model calls and more latency per decision.
- Decision: the checker comes from a different lab than the author. Why: self-review shares the author’s blind spots. Trade-off: it needs a second provider connected; with one provider, the seat falls back to the same vendor.
- Decision: a deterministic check sits in front of any model judgment. Why: a quote either appears in the source or it doesn’t, and no model should vote on that. Trade-off: faithful quotes damaged by OCR need special handling to survive.
- Decision: humans get the split vote. Why: automate the agreement, escalate the disagreement. Trade-off: a human stays in the loop on every split.
The hard part
Turning “a different lab checks the work” from a preference into a guarantee, and keeping the council standing when one vendor fails mid-run.
The first version of Concord’s model selection only applied a score penalty to the analyst’s own provider, so a strong same-vendor model could still win the challenger seat. The fix excludes the analyst’s provider outright whenever any other provider can fill the seat. Six tests were written to break the penalty-only version, and they hold across every provider pairing in the model catalog.
Separately, a live re-run lost all eleven rebuttals, the committee and the memo when one vendor’s credit ran out mid-run, and the risk register came back empty. A failed node returned a placeholder string, and the pipeline carried it forward as if it were analysis. The fix is a guard: a rebuttal replaces a desk’s position only if it actually completed.
Results
Two small benchmarks, reported with their caveats.
- Warrant. Two models from the same lab independently mapped 94 government-form fields. They agreed on 87. All 3 known errors sat inside the 7 disagreements. The 87 agreements were never audited, and one of the errors was found by inspecting a disagreement, so the setup favors that result.
- Concord. On 20 planted citations, a deterministic checker and an LLM judge agreed on 19, and 2 of those agreements were wrong: their errors overlapped instead of cancelling out. The same judge gave different verdicts on 2 of 10 citations when asked the identical question five times.
What the data supports so far is the third of my working theories: agreement is not proof. The other two, that adjudication beats synthesis and that the model that writes shouldn’t check its own work, are built into the code but not yet measured. I’m not claiming results I haven’t run.
What I’d do next
- Rerun Concord’s citation corpus and Warrant’s 94 fields with two or three vendors, and measure P(error | models agree) against P(error | models disagree), and self-check against cross-lab check.
- Score AccountWard’s vote patterns (three of three, two to one, single vote) against its synthetic corpus labels.
- Upgrade the position paper, Disagreement as Signal, to a preprint once those experiments run. Its evaluation section is a proposal until then.
Links
- Where it runs: AccountWard · Concord · Warrant + Forge · Redl
- The paper: Disagreement as Signal (position paper)
- Code: spread across the builds above; walkthrough on request.
Verified numbers
| Metric | Value | Source |
|---|---|---|
| Warrant field mappings where two models agree | 87 of 94 | mit37/warrant/docs/BENCHMARKS.md:90-119 |
| Known mapping errors that sat inside the 7 disagreements | 3 of 3 | mit37/warrant/docs/BENCHMARKS.md:90-119 |
| Citations where a deterministic checker and an LLM judge agreed, yet both were wrong | 2 of 19 | local: concord/bench/uncorrelated-result.json |
| Citations where the same LLM judge changed its verdict across 5 identical asks | 2 of 10 | local: concord/bench/stability-result.json |
| Challenger seats run on a different vendor from their analyst (live run, synthetic data room) | 11 of 11 | local: concord/scripts/live-result.json |
| Tests guarding challenger independence | 6 | local: concord/src/lib/independence.test.ts |