Work / AI Systems
Chispen
Compiles a written process into a playbook that runs, measures itself and proposes its own fixes.
- React 18
- TypeScript (strict)
- Vite 5
- React Router 6
- Anthropic SDK (streaming, optional)
- Hand-built SVG charts
TL;DR
- Most processes live in documents nobody can watch run. Chispen compiles them into software that runs, measures itself and proposes its own improvements.
- One loop: compile a process into a playbook, run it with SLAs, escalation and quorum gates, then learn from the event log and propose versioned fixes.
- The governing rule is glass-box: models and rules compile and propose, a human accepts every change, and the version history records who made it and why.
The problem
A process written in a document, like vendor onboarding or incident response, has owners, deadlines, approvals and dependencies, but nobody can see where it stalls. Improving it means someone rereads the document, guesses at the bottleneck and edits the text, with no record of what changed or whether it helped.
What I built
I architected and specified Chispen, then directed Claude Code through the build as multi-agent workflows: a product-design judge panel first, then build, review and fix passes. It runs entirely in the browser on a seeded demo workspace.
- Compile. Describe a process in plain language, or pick one of the workspace’s source documents, and the compiler streams back an executable playbook: phases, steps, owner roles, SLAs, a dependency graph and quorum approval gates, with each step carrying a short quote from the source section it came from. With an API key it calls Claude under a strict JSON contract. Without one it plays a scripted preview, and a failure in live mode falls back to that preview mid-stream, so a demo can’t dead-end.
- Run. Steps unlock when their dependencies finish. Starting a step sets its SLA deadline; a sweep flags breaches and, 12 hours past a breach, escalates the step and reassigns it. Gates count quorum votes, each with a recorded rationale. Every run action lands in an append-only event log.
- Copilot. An in-run assistant answers from a live summary of the run and the playbook’s sources, and keeps only citations that point at a section it was actually given.
- Learn. Analytics are computed from run history at render time: per-step median and p90 durations, breaches, blocks, a heat score, hotspots and cycle-time trends. A rule-based generator turns the hottest step into an improvement proposal, written as a structured diff with the run evidence attached.
- Forecast. A 500-sample Monte Carlo simulation draws lognormal step durations around each step’s measured median, schedules them in topological order and traces the critical path in every sample. It returns P50 and P90 cycle times plus how often each step sat on the critical path, which shows where a fix would pay off.
Key decisions
- Decision: glass-box versioning. A human accepts every change, and each version records its origin. Why: a process of record can’t change silently. Trade-off: nothing improves on its own; someone has to review each proposal.
- Decision: analytics derived from the event log, never stored. Why: the numbers can’t drift from what actually happened. Trade-off: recomputed on every render, which is fine at demo scale and needs a backend at production scale.
- Decision: live-to-scripted fallback. Why: an API or parse failure mid-demo shouldn’t strand the user. Trade-off: the scripted path has to be kept in step with the live contract.
The hard part
Letting proposals rewrite a live process without breaking the runs already using it. Accepted proposals and human edits change the process of record while runs are still executing against earlier versions.
Chispen never edits a version in place. Accepting a proposal deep-copies the active version’s steps and phases, applies the structured diff and mints a new semantic version that records where it came from; each run keeps its own version. The diff applier also has to keep the dependency graph valid. A merge inserts the merged step where the first original sat, deletes the originals and rewires every downstream dependency onto the merged step, deduplicated. A removal prunes the dangling edges. Downstream, the simulator’s topological sort falls back to phase order if a cycle ever appears, so a malformed graph can’t hang it.
A second case: a rejected quorum gate must not dead-end a run. A rejection blocks the step with a named reason, and resolving the blocker clears only the rejecting votes, so the gate can reach quorum again.
Results
- 12,222 lines across 38 TypeScript files and one stylesheet, with 13 routes, a pan-and-zoom flow map, a Gantt run timeline, a command palette and light and dark themes.
- A forecast engine that reports P50, P90 and per-step criticality from 500 seeded samples.
- The blueprint lays out the production path: an event-sourced backend on Postgres with the event log as the source of truth, SLA checks as scheduled evaluators that reuse the same logic, and a nightly batch in which a model files proposals through the same diff schema.
What I’d do next
- Build that backend, so runs and events persist beyond one browser.
- Check provenance quotes against the source text instead of trusting the model to copy them exactly.
- Validate the dependency graph at compile time, not only in the simulator.
Links
- Code: private, walkthrough on request.
Verified numbers
| Metric | Value | Source |
|---|---|---|
| Lines of React, TypeScript and CSS, running entirely in the browser | 12,222 | mit37/chispen/src/ |
| Monte Carlo samples per cycle-time forecast | 500 | mit37/chispen/src/domain/simulate.ts:86 |
| Run-event types in the append-only log | 13 | mit37/chispen/src/domain/types.ts:153-156 |
| Escalation after an SLA breach | 12 h | mit37/chispen/src/domain/store.ts:157-164 |