Work / AI Systems

NEXUS

Solo-built wargame adjudicator: five Claude agents debate while a 10,000-run Monte Carlo cascade does the math.

Status
archived
Role
Solo: architected the engine and the Android app, directed AI coding agents (Cursor, Antigravity)
Team
Solo entry · SCSP AI+ Expo Hackathon, Wargaming track, San Francisco round
Timeline
Apr 2026 – Apr 2026
  • Python 3.11
  • FastAPI
  • Server-Sent Events
  • NetworkX
  • Anthropic API
  • Next.js 15
  • TypeScript
  • Tailwind CSS
  • Three.js
  • Mapbox
  • Expo SDK 54
  • React Native

TL;DR

  • Built solo in one weekend for the Wargaming track of SCSP’s AI+ Expo Hackathon (San Francisco round, April 25–26, 2026), where teams could field up to five people.
  • A deterministic cascade model runs 10,000 Monte Carlo trials per scenario while five Claude agents argue over the results. The language models never touch the arithmetic.
  • Built a web engine and a React Native Android companion in the same weekend, and demoed NEXUS live to the Wargaming track’s national-security judging panel, who probed the simulation logic and the AI adversary.

The problem

The track asked teams to use AI to adjudicate wargames: decide what happens when a move lands, generate scenarios and play an adaptive opponent. A language model asked “what happens if this node fails?” will give a confident, fluent answer with no way to tell whether it is right. Supply chains fail through dependencies, and dependency cascades are a math problem.

What I built

My design rule was to keep the language models away from the arithmetic.

  • Cascade model. A Python layer models a supply chain as a NetworkX directed graph: a seed asset, intermediate processors and a terminal asset, with a dependency weight on every edge and a resilience threshold on every node. A node fails once the weight of its failed inputs crosses its threshold.
  • Monte Carlo. The engine runs that cascade 10,000 times with jittered thresholds and reports the failure probability with a 95% Wald confidence interval. Weighted betweenness centrality ranks the chokepoints.
  • Council. Five Claude agents take turns over a shared transcript, each answering the ones before it: an Economist, an Infrastructure Architect, a Logistics Chief, a Risk Analyst and an Executive Arbiter. The simulation runs between their turns.
  • Interface. Results stream over Server-Sent Events to a Next.js terminal UI with a hand-built Three.js globe of 128 infrastructure nodes, Mapbox satellite views, a multi-target queue and a “harden this node” counterfactual slider.
  • Android companion. An Expo/React Native app with a Three.js globe of all 128 nodes, a node browser with on-demand AI briefings you can export to PDF, and a command console. I compiled it to a release APK and ran it on my own phone.

I architected both pieces and directed AI coding agents through the build: Cursor for the engine, Google Antigravity for the app.

Key decisions

  • Decision: deterministic math, language models for argument only. Why: the probability has to be checkable; the debate adds context, not numbers. Trade-off: the agents can only reason about what the model computes.
  • Decision: Monte Carlo with a reported confidence interval. Why: a single cascade run is an anecdote; 10,000 jittered runs give a probability and its uncertainty. Trade-off: each scenario costs 10,000 graph traversals.
  • Decision: stream every step over Server-Sent Events. Why: judges watch the council argue in real time instead of waiting on a spinner. Trade-off: more moving parts to demo live.

The hard part

A simulation that always says “catastrophe” is broken, not alarming.

The first version came back catastrophic for every scenario. I flagged that the verdict never changed. The cause was degenerate math: edge weights exceeded node thresholds, so a single upstream failure reached the terminal asset in almost every trial (0.958 on the first smoke test). The five-agent council was arguing confidently over a number that carried no information.

I had the model rebuilt with hand-set weight and threshold profiles per category, threshold jitter widened to ±20% so thresholds actually vary across trials, and a recalibrated priority scale. After that, the outputs discriminated between targets. One exported run reported a 10.51% cascade probability: 1,051 of 10,000 trials, with a 95% confidence interval of 9.91–11.11%. The interval checks out by hand: 0.1051 ± 1.96 × √(0.1051 × 0.8949 / 10,000) ≈ 0.1051 ± 0.0060.

Results

  • A working engine and a release-built Android companion, both built in one weekend by a solo entrant in a track where teams could have five.
  • Outputs that discriminate between targets after the repair, each with its confidence interval.
  • Demoed live to the Wargaming track’s national-security judging panel, who probed the simulation logic and the AI adversary.

What I’d do next

  • Calibrate the hand-set weight and threshold profiles against real supply-chain disruption data.
  • Replace the sequential debate with adjudication: agents propose, a checker from another lab challenges, and dissent stays on the record. That is where The Council went next.
  • Publish a cleaned copy of the engine.
  • Code: private, walkthrough on request.
  • What came next: The Council

Verified numbers

MetricValueSource
Monte Carlo trials per scenario run10,000local: NEXUS build log (Cursor export), line 3066
Cascade probability in one exported run, with its 95% Wald CI10.51% (9.91–11.11%)local: NEXUS-Brief-NX-2026-0426-6FQR.pdf
Infrastructure nodes on the globe (23 countries plus international chokepoints)128mit37/NEXUS/src/data/targetNodes.ts
Council agents taking turns over a shared transcript5local: NEXUS build log (Cursor export), lines 3024-3040
Android release build (arm64 + armv7, Android 7.0+)74.4 MB APKlocal: app-release.apk
m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h