Work / AI Systems
NEXUS
Solo-built wargame adjudicator: five Claude agents debate while a 10,000-run Monte Carlo cascade does the math.
- Python 3.11
- FastAPI
- Server-Sent Events
- NetworkX
- Anthropic API
- Next.js 15
- TypeScript
- Tailwind CSS
- Three.js
- Mapbox
- Expo SDK 54
- React Native
TL;DR
- Built solo in one weekend for the Wargaming track of SCSP’s AI+ Expo Hackathon (San Francisco round, April 25–26, 2026), where teams could field up to five people.
- A deterministic cascade model runs 10,000 Monte Carlo trials per scenario while five Claude agents argue over the results. The language models never touch the arithmetic.
- Built a web engine and a React Native Android companion in the same weekend, and demoed NEXUS live to the Wargaming track’s national-security judging panel, who probed the simulation logic and the AI adversary.
The problem
The track asked teams to use AI to adjudicate wargames: decide what happens when a move lands, generate scenarios and play an adaptive opponent. A language model asked “what happens if this node fails?” will give a confident, fluent answer with no way to tell whether it is right. Supply chains fail through dependencies, and dependency cascades are a math problem.
What I built
My design rule was to keep the language models away from the arithmetic.
- Cascade model. A Python layer models a supply chain as a NetworkX directed graph: a seed asset, intermediate processors and a terminal asset, with a dependency weight on every edge and a resilience threshold on every node. A node fails once the weight of its failed inputs crosses its threshold.
- Monte Carlo. The engine runs that cascade 10,000 times with jittered thresholds and reports the failure probability with a 95% Wald confidence interval. Weighted betweenness centrality ranks the chokepoints.
- Council. Five Claude agents take turns over a shared transcript, each answering the ones before it: an Economist, an Infrastructure Architect, a Logistics Chief, a Risk Analyst and an Executive Arbiter. The simulation runs between their turns.
- Interface. Results stream over Server-Sent Events to a Next.js terminal UI with a hand-built Three.js globe of 128 infrastructure nodes, Mapbox satellite views, a multi-target queue and a “harden this node” counterfactual slider.
- Android companion. An Expo/React Native app with a Three.js globe of all 128 nodes, a node browser with on-demand AI briefings you can export to PDF, and a command console. I compiled it to a release APK and ran it on my own phone.
I architected both pieces and directed AI coding agents through the build: Cursor for the engine, Google Antigravity for the app.
Key decisions
- Decision: deterministic math, language models for argument only. Why: the probability has to be checkable; the debate adds context, not numbers. Trade-off: the agents can only reason about what the model computes.
- Decision: Monte Carlo with a reported confidence interval. Why: a single cascade run is an anecdote; 10,000 jittered runs give a probability and its uncertainty. Trade-off: each scenario costs 10,000 graph traversals.
- Decision: stream every step over Server-Sent Events. Why: judges watch the council argue in real time instead of waiting on a spinner. Trade-off: more moving parts to demo live.
The hard part
A simulation that always says “catastrophe” is broken, not alarming.
The first version came back catastrophic for every scenario. I flagged that the verdict never changed. The cause was degenerate math: edge weights exceeded node thresholds, so a single upstream failure reached the terminal asset in almost every trial (0.958 on the first smoke test). The five-agent council was arguing confidently over a number that carried no information.
I had the model rebuilt with hand-set weight and threshold profiles per category, threshold jitter widened to ±20% so thresholds actually vary across trials, and a recalibrated priority scale. After that, the outputs discriminated between targets. One exported run reported a 10.51% cascade probability: 1,051 of 10,000 trials, with a 95% confidence interval of 9.91–11.11%. The interval checks out by hand: 0.1051 ± 1.96 × √(0.1051 × 0.8949 / 10,000) ≈ 0.1051 ± 0.0060.
Results
- A working engine and a release-built Android companion, both built in one weekend by a solo entrant in a track where teams could have five.
- Outputs that discriminate between targets after the repair, each with its confidence interval.
- Demoed live to the Wargaming track’s national-security judging panel, who probed the simulation logic and the AI adversary.
What I’d do next
- Calibrate the hand-set weight and threshold profiles against real supply-chain disruption data.
- Replace the sequential debate with adjudication: agents propose, a checker from another lab challenges, and dissent stays on the record. That is where The Council went next.
- Publish a cleaned copy of the engine.
Links
- Code: private, walkthrough on request.
- What came next: The Council
Verified numbers
| Metric | Value | Source |
|---|---|---|
| Monte Carlo trials per scenario run | 10,000 | local: NEXUS build log (Cursor export), line 3066 |
| Cascade probability in one exported run, with its 95% Wald CI | 10.51% (9.91–11.11%) | local: NEXUS-Brief-NX-2026-0426-6FQR.pdf |
| Infrastructure nodes on the globe (23 countries plus international chokepoints) | 128 | mit37/NEXUS/src/data/targetNodes.ts |
| Council agents taking turns over a shared transcript | 5 | local: NEXUS build log (Cursor export), lines 3024-3040 |
| Android release build (arm64 + armv7, Android 7.0+) | 74.4 MB APK | local: app-release.apk |