Work / AI Systems

The Council

Agents do the work, other models check it. Built nine ways, then measured, with the negative results reported.

Status
research
Role
Architect: specified nine builds, directed Claude Code, Codex and Cursor agents, designed the benchmarks
Timeline
Apr 2026 – Aug 2026
  • TypeScript
  • Rust
  • Python
  • Anthropic API
  • OpenAI-compatible APIs
  • Vitest

Live demo

Council playground

Several answers to one question. The disagreement layer finds the claims they split on, and the threshold decides what a human has to look at.

QuestionA $1,000 deposit earns 5% simple interest a year. What is the balance after 3 years?

A

Simple interest pays 5% of the original $1,000 each year: $50 × 3 = $150. The balance is $1,150.00.

B

Interest accrues on the principal only, with no compounding, so the deposit earns $50 a year. After three years it stands at $1,150.

Cdissents on 2

Compounding annually: $1,000 × 1.05³ = $1,157.63.

D

Three years of $50 in interest is $150, for a balance of $1,150.00.

Worst ∆0.33

2 of 2 claims go to a human.

ClaimAnswers, groupedAgreementDecision
Balance after 3 years1,150 A B D1,157.63 C75% ∆ 0.25human majority is right
Interest methodsimple A Bcompound C67% ∆ 0.33human majority is right

One answer compounds when the question says simple interest. The other three agree, and the split still reaches a human at the default threshold.

Scripted examples written for this demo. They are not output from any AI model, and the labels A to E are placeholders.

TL;DR

  • Between April and August 2026 I built the same idea nine times: agents that do the work, and a council of other models that checks it.
  • The design moved from synthesis (merge everyone’s answer) to adjudication (keep dissent on the record, challenge each analyst with another vendor’s model, check every citation deterministically, send split votes to a human).
  • Then I measured the premise. Where the data cut against it, I report that too: agreement is evidence, not proof.

The problem

A single model is most dangerous when it is confidently wrong, and asking it to check its own work shares its blind spots. The common fix, asking several models and merging their answers, averages the errors in instead of surfacing them. I wanted an architecture where disagreement is treated as a signal to act on.

What I built

I specified each build and directed AI coding agents (Claude Code, Codex and Cursor) to write the code. The lineage:

#BuildWhenProtocol
1NEXUS, wargame adjudicationApr 2026Sequential agent debate, then an arbiter’s synthesis, with a Monte Carlo check
2Cross-vendor M&A “grand council”May 2026Per-vendor agent swarms, then cross-vendor arbitration and a red-team loop
3Coding-agent review councilJun 2026Five reviewers, then a judge’s synthesis
4BlueCollarPal takeoffJul 2026Estimator against skeptic; a deterministic reconciler sends challenged lines to a human
5Redl Council paneJul 2026Layered synthesis pyramid, 4 → 3 → 1
6QuorumJul 2026Redl’s pyramid with diligence roles
7ConcordJul–Aug 2026Cross-examination by a different-vendor challenger, deterministic tie-out, evidence auditor, mandatory dissent
8Warrant cross-checkJul 2026Two independent mapping runs, with human adjudication stored as data
9AccountWard classifierJul–Aug 2026Three-method vote; below 0.9 confidence goes to a human

The reference build is Concord’s seven-stage diligence pipeline. Parallel analyst desks run first. A cross-examination round seats each challenger on a different provider from its analyst whenever a second provider is connected. Then come a devil’s advocate, a deterministic citation tie-out with no model in the loop, an evidence auditor, a risk committee that records disputes with both positions, and an investment-committee memo with a mandatory dissent section.

Key decisions

  • Decision: adjudicate, don’t synthesize. Why: merging answers averages errors in; picking and justifying keeps the dissent visible. Trade-off: more model calls and more latency per decision.
  • Decision: the checker comes from a different lab than the author. Why: self-review shares the author’s blind spots. Trade-off: it needs a second provider connected; with one provider, the seat falls back to the same vendor.
  • Decision: a deterministic check sits in front of any model judgment. Why: a quote either appears in the source or it doesn’t, and no model should vote on that. Trade-off: faithful quotes damaged by OCR need special handling to survive.
  • Decision: humans get the split vote. Why: automate the agreement, escalate the disagreement. Trade-off: a human stays in the loop on every split.

The hard part

Turning “a different lab checks the work” from a preference into a guarantee, and keeping the council standing when one vendor fails mid-run.

The first version of Concord’s model selection only applied a score penalty to the analyst’s own provider, so a strong same-vendor model could still win the challenger seat. The fix excludes the analyst’s provider outright whenever any other provider can fill the seat. Six tests were written to break the penalty-only version, and they hold across every provider pairing in the model catalog.

Separately, a live re-run lost all eleven rebuttals, the committee and the memo when one vendor’s credit ran out mid-run, and the risk register came back empty. A failed node returned a placeholder string, and the pipeline carried it forward as if it were analysis. The fix is a guard: a rebuttal replaces a desk’s position only if it actually completed.

Results

Two small benchmarks, reported with their caveats.

  • Warrant. Two models from the same lab independently mapped 94 government-form fields. They agreed on 87. All 3 known errors sat inside the 7 disagreements. The 87 agreements were never audited, and one of the errors was found by inspecting a disagreement, so the setup favors that result.
  • Concord. On 20 planted citations, a deterministic checker and an LLM judge agreed on 19, and 2 of those agreements were wrong: their errors overlapped instead of cancelling out. The same judge gave different verdicts on 2 of 10 citations when asked the identical question five times.

What the data supports so far is the third of my working theories: agreement is not proof. The other two, that adjudication beats synthesis and that the model that writes shouldn’t check its own work, are built into the code but not yet measured. I’m not claiming results I haven’t run.

What I’d do next

  • Rerun Concord’s citation corpus and Warrant’s 94 fields with two or three vendors, and measure P(error | models agree) against P(error | models disagree), and self-check against cross-lab check.
  • Score AccountWard’s vote patterns (three of three, two to one, single vote) against its synthetic corpus labels.
  • Upgrade the position paper, Disagreement as Signal, to a preprint once those experiments run. Its evaluation section is a proposal until then.

Verified numbers

MetricValueSource
Warrant field mappings where two models agree87 of 94mit37/warrant/docs/BENCHMARKS.md:90-119
Known mapping errors that sat inside the 7 disagreements3 of 3mit37/warrant/docs/BENCHMARKS.md:90-119
Citations where a deterministic checker and an LLM judge agreed, yet both were wrong2 of 19local: concord/bench/uncorrelated-result.json
Citations where the same LLM judge changed its verdict across 5 identical asks2 of 10local: concord/bench/stability-result.json
Challenger seats run on a different vendor from their analyst (live run, synthetic data room)11 of 11local: concord/scripts/live-result.json
Tests guarding challenger independence6local: concord/src/lib/independence.test.ts
m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h