Work / Fintech & RegTech

Warrant + Forge

Estate-settlement engine where no fact reaches a decision without a verbatim quote found in a source document.

Status
built
Role
Architected Warrant and directed Claude Code through it; built the Warrant–Forge merge, the executor portal and the Android app; presented the demo
Collaboration
Built with Fluxxion88
Timeline
Jul 2026 – Jul 2026
  • TypeScript
  • React 19
  • Vite 7
  • Tailwind CSS 4
  • Rust
  • Tauri 2
  • Python (pypdf, pdfplumber)
  • Kotlin (Android WebView)
  • Vitest
  • Anthropic, OpenAI, Gemini, DeepSeek and Moonshot APIs

TL;DR

  • Built at an estate-settlement AI hackathon in San Francisco (July 27–28, 2026) with my teammate, Fluxxion88. I architected Warrant, the rules and verification engine, and directed Claude Code through 42 commits in about 24 hours; my teammate built Forge, the form compiler.
  • Models may propose facts but not assert them. Every fact needs a warrant: a verbatim quote that a deterministic matcher finds in a held document, a recorded derivation, or a record path.
  • On day two I merged the two engines in my own build, cross-checked their form pipelines against each other, built an executor portal as a Kotlin Android app, and presented the demo.

The problem

Settling an estate means turning a pile of documents into decisions: which assets exist, which statutes apply, which forms to file. Each decision rests on a fact, like a date-of-death balance or a filing fee. A language model will read the pile and state those facts fluently, including ones that appear nowhere in it. In estate work the number is the claim, and an invented one ends up in a government form.

What I built

I didn’t start from zero. At 13:09 on day one I forked my own Concord diligence engine and kept its quote verifier, Rust document ingestion and multi-provider model layer. Everything estate-specific was built at the event.

  • Fact ledger. Three kinds of warrant: quote, derivation and record. Only verified facts reach the rules; anything that fails is quarantined, visible to a human and invisible to every rule.
  • Rules as data. Each rule carries its citation and effective date and is evaluated in three-valued logic, so a missing fact blocks a decision and names what it needs instead of guessing. Probate packs for California, New York, Pennsylvania and Texas, including cited California thresholds under Probate Code §§ 13100 and 13151.
  • County registry. All 58 California counties, with the 55 not yet researched flagged as such instead of guessed at.
  • Reactor. Re-evaluates only the decisions a changed fact touches.
  • Asset discovery. Read 157 transactions on a synthetic estate’s bank statement, found 12 recurring charges, and surfaced 3 assets mentioned in no document, kept as hypotheses, never as facts.
  • Approval gate. Scored on reversibility × blast radius, a mentor’s framing that I turned into code.
  • Form onboarding. A model maps each PDF form field and has to cite the label printed beside it, checked against the page geometry.

On day two I merged Warrant with Forge, built by my teammate. Forge uses a model once per form at build time, has a human approve the binding, then fills forms deterministically with zero model calls. The team’s thesis: we don’t fill forms with AI, we use AI to build form-fillers. Then I turned the engine outward into a portal for the executor and shipped it as an Android app about 15 minutes before the presentations.

Key decisions

  • Decision: facts need warrants; models only propose. Why: a fluent invented number is the failure that matters here. Trade-off: some true facts get quarantined until a human supplies the source.
  • Decision: three-valued logic for rules. Why: “unknown” has to block a decision, not default to false. Trade-off: plans come back with explicit unresolved lists instead of a clean answer.
  • Decision: disagreements go to a human worklist, never to a vote. Why: in the benchmark, every known-wrong mapping sat inside the disagreements. Trade-off: a person reviews every split.
  • Decision: fork my own verifier instead of writing a new one at the event. Why: the verifier, ingestion and provider layers already existed, which left the event for the estate logic. Trade-off: the codebase inherited modules the estate domain doesn’t use.

The hard part

A number-substitution hole in the verifier. N-gram coverage has to be tolerant so that a faithful quote survives reflowed lines and smart quotes. Change one figure in an eleven-token sentence, though, and only one of eight 4-gram windows misses: coverage is 0.875, which clears the 0.85 bar and grades “verified”.

The hole turned up when a test written to assert that an invented filing fee gets refused didn’t refuse it. The fix requires every digit-bearing token in a quote to appear in the document, by substring, so $435 quoted from a page that prints $435.00 still passes. Leniency survives only in the direction that can’t fabricate a value. All 41 facts from the recorded live run still verified after the change, and a corruption sweep caught 24 of 24 number substitutions and 41 of 41 wholesale rewrites.

It caught 0 of 41 cases where a genuine quote carried a wrong value. I published that boundary in the benchmark instead of hiding it.

Results

  • 41 of 41 proposed facts verified on a recorded live extraction over 9 synthetic documents.
  • Two models mapping the same four government forms agreed on 87 of 94 shared fields, and all 3 known-wrong mappings sat inside the 7 disagreements.
  • In my merge, cross-checking the two independently built form pipelines caught real bugs on both sides: Forge’s draft Form 8821 was writing the signature date into the title box, and Warrant’s SS-4 used an SSN where the schema wanted the responsible party’s TIN. After the fixes, the two pipelines agreed on every comparable field across all four forms, and pixel-diffing every written field confirmed 255 of 255 across 14 PDFs.
  • Every benchmark metric is tagged with its provenance: measured, derived, assumed or tainted.

What I’d do next

  • Close the value-only gap: check a quoted value against the document’s own structure, not just its text.
  • Research the remaining California counties instead of flagging them.
  • Fold the engine into AccountWard, which grew out of this build, as one platform for professional fiduciaries.

Verified numbers

MetricValueSource
Proposed facts verified in a recorded live extraction over 9 synthetic documents41 of 41mit37/warrant/docs/BENCHMARKS.md
Number substitutions caught by the quote verifier24 of 24mit37/warrant/docs/BENCHMARKS.md:154-159
Genuine quotes carrying a wrong value caught (a published limit)0 of 41mit37/warrant/docs/BENCHMARKS.md:186
Form fields where two models agreed; all 3 known errors sat in the 7 disagreements87 of 94mit37/warrant/docs/BENCHMARKS.md:90-119
Written form fields pixel-verified after the Warrant–Forge cross-check, across 14 PDFs255 of 255local: warrant-backup/warrant-plus-final.bundle
TypeScript test cases (static count) plus 11 Rust tests536mit37/warrant/src/**/*.test.ts
m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h