Work / AI Systems

ForgeCouncil

The council as an AI engineering team: reviewer agents plus a judge, in a safe-by-default CLI and desktop app.

Status
explored
Role
Solo: specified and directed Claude Code
Timeline
Jun 2026 →
  • TypeScript
  • Node.js
  • Commander.js
  • Zod
  • Vitest
  • Electron
  • React
  • Vite

TL;DR

  • Prototyped the council as an AI engineering team in one Claude Code session: five reviewer agents and a judge that returns a schema-validated verdict.
  • Safe by default: it writes only inside its own folder unless you pass --apply, runs only package scripts, and never reads .env.
  • The mock run exposed the flaw that shaped every council I built after it: one summarizing judge inherits every reviewer’s blind spot.

What I built

A TypeScript CLI with a local server and an Electron app. Product, architecture, bug-hunter, security and QA agents review a repository. A Council Judge then returns a Zod-validated verdict (proceed, revise, needs-human or reject) with a confidence, a risk level and questions for a person. The server streams progress over WebSockets, and the desktop app packaged as a Windows installer.

Safe by default → an agent council pointed at a real repo must not be able to damage it → it can fix less on its own. Everything lands under .forge/ unless you pass --apply. Even then it only auto-creates docs and example files, runs only npm or pnpm scripts, and never opens .env.

What I learned

On the default mock provider, the judge returned PROCEED from template findings that missed a hardcoded-key TODO the repo scanner had already flagged. A judge that summarizes can’t catch what its reviewers missed. My later councils, starting with Concord, replaced the summarizing judge with cross-vendor challengers and deterministic checks.

Status

Explored. The judge ran on one provider and never against a live model, and the OpenAI and Ollama providers were only scaffolded. The code sits on a drive that isn’t mounted right now, so this page is written from the build session.

m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h