Work / AI Systems

AI-assisted game production pipeline

Generating and play-testing genre games on Roblox with coding agents: three genres in parallel, one routed to verify.

Status
explored
Role
Solo: briefs, model routing and orchestration of AI coding agents
Timeline
Sep 2026 →
  • Freebuff agent threads
  • GLM-5.3-flash
  • DeepSeek-v4-flash
  • Solar Pro 4
  • GPT-5.6-luna (verification pass)
  • Claude Code
  • Luau
  • Rojo

TL;DR

  • I wanted to know how far coding agents can take a Roblox game from a written brief.
  • I ran three agent threads in parallel on free-tier models, one genre each, and routed a stronger model to verify.
  • Generation turned out to be the cheap part. Getting a build to play correctly in Studio is the real work.

The run

Three Freebuff threads, one per genre: an extraction-horror shooter; THE TRIALS, a 24-player survival game; and COLLAPSE, a 24-player falling-floor party game. I pushed all three with the same finish-it prompt within seconds of each other, twice, and routed the shooter to GPT-5.6-luna for a fine-tooth-comb verification pass. Each came back as a Rojo-structured Luau codebase with design docs and a built place file.

When the free sessions ran out, I moved the shooter into a multi-agent rebuild in Claude Code, which became BLACKSITE: NULL. That’s four codebases from one run.

What I concluded

Generation is the cheap part. Getting a build to play correctly in Studio is the real work, and so far only BLACKSITE: NULL has been through a Studio playtest.

From the same run

  • COLLAPSE: a heap-scheduled tile engine where idle tiles cost nothing per tick.
  • THE TRIALS: server-validated combat and a headless two-match smoke harness.
m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h