Work / AI Systems
ForgeCouncil
The council as an AI engineering team: reviewer agents plus a judge, in a safe-by-default CLI and desktop app.
- TypeScript
- Node.js
- Commander.js
- Zod
- Vitest
- Electron
- React
- Vite
TL;DR
- Prototyped the council as an AI engineering team in one Claude Code session: five reviewer agents and a judge that returns a schema-validated verdict.
- Safe by default: it writes only inside its own folder unless you pass
--apply, runs only package scripts, and never reads.env. - The mock run exposed the flaw that shaped every council I built after it: one summarizing judge inherits every reviewer’s blind spot.
What I built
A TypeScript CLI with a local server and an Electron app. Product, architecture, bug-hunter, security and QA agents review a repository. A Council Judge then returns a Zod-validated verdict (proceed, revise, needs-human or reject) with a confidence, a risk level and questions for a person. The server streams progress over WebSockets, and the desktop app packaged as a Windows installer.
Safe by default → an agent council pointed at a real repo must not be able to damage it → it can fix less on its own. Everything lands under .forge/ unless you pass --apply. Even then it only auto-creates docs and example files, runs only npm or pnpm scripts, and never opens .env.
What I learned
On the default mock provider, the judge returned PROCEED from template findings that missed a hardcoded-key TODO the repo scanner had already flagged. A judge that summarizes can’t catch what its reviewers missed. My later councils, starting with Concord, replaced the summarizing judge with cross-vendor challengers and deterministic checks.
Status
Explored. The judge ran on one provider and never against a live model, and the OpenAI and Ollama providers were only scaffolded. The code sits on a drive that isn’t mounted right now, so this page is written from the build session.