Work / Fintech & RegTech
Concord
Multi-vendor M&A diligence engine: 11 agent desks, cross-lab challengers and a model-free tie-out on every quote.
- Tauri 2
- Rust (pdf-extract, calamine, quick-xml)
- React 19
- TypeScript
- Vite 7
- Tailwind CSS 4
- SQLite
- Anthropic and OpenAI-compatible APIs
- Vitest
TL;DR
- I built Concord to staff a diligence desk with agents instead of asking one model for an opinion: a Tauri desktop app that runs a seven-stage engagement over a data room, named the way a deal team names it.
- Every finding has to cite a verbatim quote, and a deterministic tie-out with no model in the loop grades each quote against the document it cites.
- One live engagement returned 152 cited findings in 50 minutes for $4.995; all 11 challenges ran on a different vendor’s model from the analyst they checked.
The problem
Asking one model to review a data room gives you one opinion, unverified, from a system that can’t tell a quote it read from a quote it invented. Real diligence runs in parallel workstreams, gets challenged by people who didn’t do the work, and ties every claim back to a document. I wanted the software to work the same way.
What I built
I architected and specified Concord, and directed AI coding agents (Claude Code) to write it. It ingests PDFs, spreadsheets, Word files and decks, then runs seven stages:
- Data Room Review. One long-context model builds a fact ledger.
- Workstream Diligence. Eleven desks run in parallel, each with its own charter and red-flag library.
- Cross-Examination. Every desk faces a challenger model routed away from the analyst’s own AI lab.
- Devil’s Advocate Review.
- Source Verification. The deterministic tie-out grades every quote as verified, loose or unsupported.
- Risk Committee. Disputes are recorded with both positions.
- IC memo. PROCEED, CONDITIONS or DECLINE.
Around the engine I built the controls a real engagement needs:
- counterparty documents fenced as untrusted input and screened for prompt injection;
- per-provider rate limits with a hard dollar cap;
- token budgets that name the documents they had to leave out;
- resumable checkpoints;
- a consolidator that merges duplicate findings, can only raise severity and can only lower verification.
Key decisions
- Decision: a deterministic tie-out, not a model, decides whether a quote exists. Why: a quote either appears in the cited document or it doesn’t, and a model shouldn’t vote on that; the benchmark I built later showed an LLM judge, asked the same question five times, calling a fabricated paraphrase “present” three times. Trade-off: quotes damaged by OCR need a separate fold to survive, and that fold is capped at “loose”.
- Decision: the challenger comes from a different vendor than the analyst. Why: a reviewer from the same lab shares the author’s blind spots. Trade-off: it needs a second provider connected; with one, it falls back to the same vendor.
- Decision: the pipeline takes an injected completion function. Why: it runs offline in tests and from a Node harness for live runs. Trade-off: one more seam to keep in sync with the desktop app.
The hard part
A failure marker that looked like success. Node failures are contained rather than fatal: a failed node returns a placeholder string instead of throwing. But the placeholder was a non-empty string, and non-empty strings are truthy.
During a live re-run, one vendor’s credit ran out mid cross-examination. Every rebuttal returned the marker, each marker overwrote the desk’s completed analysis, and the markers flowed into the committee and memo prompts as if they were analysis. An engagement with eleven successful analyses behind it produced an empty register after $2.97 of spend.
The first fix guarded against empty rebuttals, and the regression test still failed; that failing test is what exposed the truthy sentinel. The real fix has three parts: one normalized failure marker with an exported check, applied wherever node output is carried forward; a provider-down error that stops calling a provider once it fails; and an outage test that pins the behavior.
The companion story: a two-desk smoke run came back empty from every analyst because adaptive thinking consumed the whole output cap. An explicit effort setting fixed it.
Results
- One live engagement over a synthetic five-document data room, driven through the engine’s Node harness: 152 cited findings in 50.4 minutes for $4.995. 100 tied out cleanly, 44 graded loose and 8 unsupported.
- A 23-case planted-fabrication benchmark for the tie-out: 15 fabrications kept out of “verified” and 8 faithful quotes kept, once I added an OCR fold.
- The live runs paid for themselves by breaking things, and both failures they exposed are fixed and tested.
Then I asked whether Concord already existed and directed a multi-agent market teardown to find out. It placed Concord in a crowded, well-funded field and concluded that only two pieces were defensible: the blocking tie-out and the per-finding cross-vendor challenge. The engine outlived the product: Concord’s verifier, ingestion and provider layers are what Warrant + Forge was forked from. Concord has been dormant since August 10, 2026.
What I’d do next
- Publish the tie-out benchmark as a public adversarial test, which the teardown named as the strongest next step.
- Replay a superseded document through the risk register in the UI; the recompute exists as a library, with zero model calls.
Links
- Code: local, walkthrough on request.
- Lineage: Redl → Quorum → Concord → Warrant + Forge → AccountWard
- The research thread: The Council
Verified numbers
| Metric | Value | Source |
|---|---|---|
| Cited findings from one live engagement on a synthetic five-document data room | 152 in 50.4 min | local: concord/scripts/raw-findings.json |
| Findings whose quote tied out cleanly (44 loose, 8 unsupported) | 100 of 152 | local: concord/scripts/raw-findings.json |
| Challenges run on a different vendor from the analyst they checked | 11 of 11 | local: com.concord.app/concord.db |
| Cost of that engagement | $4.995 | local: com.concord.app/concord.db |
| Tries in which an LLM judge called a fabricated paraphrase “present” | 3 of 5 | local: concord/bench/stability-result.json |
| Vitest cases plus 9 Rust tests | 103 | local: concord/src/lib/*.test.ts |