Lab
17 things I investigated, and what I concluded.
Explorations
Model routing in practice
I’ve routed models by task, cost and vendor in three codebases built by AI coding agents under my direction: JobApp tiers work across Claude models, Redl seats each council member on a local or cloud model, and Concord scores models on capability and price.
→ Routing by cost is table stakes. The rule worth keeping is about verification: the model that checks the work comes from a different provider than the model that did it.
AI-assisted game production pipeline
How far can coding agents take a Roblox game from a written brief? I ran three agent threads in parallel on free-tier models, one genre each, pushed each with the same finish-it prompt, and routed a stronger model to do a verification pass.
→ Generation is the cheap part: each thread came back as a Rojo-structured Luau codebase with design docs and a built place file. Getting a build to play correctly in Studio is the real work, and so far only BLACKSITE: NULL has been through a Studio playtest.
Quorum
Could M&A due diligence run entirely on open-weight models on the user’s own machine? I forked my Redl app into Quorum and directed Claude Code: five workstream leads, a risk committee with a devil’s advocate, and an IC memo, on Ollama by default.
→ Its gated end-to-end test on a local 7B model produced a PROCEED WITH CONDITIONS memo that flagged a change-of-control clause. Then I forked it again into Concord, to put frontier models from several labs on the same desk.
AI due-diligence market map
After running Concord live, I asked whether it already existed and directed a multi-agent research sweep across six segments of the AI due-diligence market, from PE-diligence software to frontier labs, scoring every Concord feature against them.
→ Multi-model routing, bring-your-own-key desktop and devil’s-advocate review came back commoditised. Two pieces survived: a blocking citation tie-out and per-finding cross-vendor challenge that keeps dissent. Its top recommendation, an adversarial benchmark, is what I built next.
Up-To
What’s the least planning it takes to get friends out the door? I explored Up-To, a plans app with one-tap invites and an “I’m down” RSVP, then a React Native availability grid and a SwiftUI iMessage extension stub, directing Google’s Antigravity agent.
→ The drag-to-paint 5 × 12 availability grid got built in one short session; the rest stops at hardcoded invites. There’s no backend yet, and the iMessage extension was never compiled. Supabase and the Xcode build were the planned next steps.
ForgeCouncil
Can a council of reviewer agents plus a judge act as an AI engineering team on a real repo? I directed Claude Code through a safe-by-default TypeScript CLI and Electron app: product, architecture, bug-hunter, security and QA reviewers, then a Zod-validated verdict.
→ A single summarizing judge is the weak seat: on the mock provider it returned PROCEED over template findings that missed a hardcoded-key TODO the repo scanner had flagged. My later councils replaced it with cross-vendor challengers and deterministic checks.
Colossus Wake
How much of a 2D action-platformer can a coding agent build in one night from written task specs? I directed OpenAI’s Codex agent through a Godot 4.6 prototype set on a sleeping titan.
→ It got the systems down fast: a tunable movement controller with coyote time, jump buffering and a dash, four enemy archetypes on a shared base class, a boss framework and three zones. No art and no export yet; parked since June 2026.
Oblivion
Can a messenger and a wallet share one app without inventing any crypto? I composed Waku end-to-end-encrypted messaging, a non-custodial Ethereum wallet and a Farcaster feed around one encrypted local vault, directing OpenAI’s Codex agent in a two-hour session.
→ Composition held up: PBKDF2-SHA256 at 310,000 iterations, AES-GCM-256 with a fresh IV on every vault write, and two local clients run side by side to test node-to-node messaging. Multi-hop routing and the native desktop build stayed on the roadmap.
ADDE: adversarial due-diligence engine
What if each AI lab ran its own agent swarm over an M&A data room, and a council with one representative per lab dropped any risk claim only one lab raised without evidence? I specified it and directed OpenAI’s Codex agent through the build.
→ The first live multi-provider run exited cleanly but wasn’t a clean result: provider errors in every department, valuation figures from only some of them, and a planted receivable trap slipped through. Arbitration between models needs deterministic checks under it.
Project Omni
Hold a hotkey, speak, and the text streams into whatever app has focus. I directed Cursor’s agent through the native Tauri 2 / Rust half of that dictation app in one four-phase session, type-checking after every phase.
→ The client came out solid: three threads, a WebSocket relay that backs off from 500 ms to 30 s and drops audio rather than block capture, and native permission probes on macOS and Windows. The transcription relay was never built, so it stays a client-side spike.
Provably-fair outcome engine
Can every random outcome be re-derived and checked by the person it affects? I specified a commit-reveal RNG and directed an AI coding agent through it: the server commits to a seed hash, and every draw is HMAC-SHA256 of server seed, client seed and nonce.
→ Any round can be replayed from its seeds, multi-draw outcomes like card shuffles pull their own HMAC sub-stream, and expected value is set in the outcome mapping itself. Balances sit in a Postgres ledger where every settlement runs in a row-locked transaction.
Offline speech-to-notes (Java)
Can meeting notes work offline first? The idea, rebuilt in Java 17: a Swing app that streams 16 kHz microphone audio in 100 ms buffers into an on-device Vosk recognizer, then asks Gemini for structured notes and answers questions about a screenshot.
→ It stays dependency-light, with JSON built and parsed by hand over java.net.http and no API key in source. It was never compiled or run, and its hard-coded model name needs updating before it will be.
MeetWise
What does a meeting copilot need beyond a transcript? MeetWise, in Electron and React 19: a live speech-to-text transcript plus screen sampling, feeding one-line suggestions into an always-on-top overlay, then notes and a follow-up email after the call.
→ The renderer compiled to a production bundle with a 1.4 s debounce before each suggestion, OpenAI-compatible or fully local Ollama back ends, and a mock fallback so the UI always runs. Live speech-to-text inside Electron is still unproven.
FitFindr
Can outfit recognition run on the phone instead of the cloud? FitFindr is a Kotlin Android app wired to run a 4-billion-parameter vision-language model on Snapdragon’s Hexagon NPU through the Nexa AI SDK.
→ The debug build compiled to a 111 MB APK, with the model pulled in-app from Hugging Face and every SDK call contained in one reflection bridge. Proving real output on a Snapdragon device is the next step.
EyeOS
Can a plain webcam turn gaze into a safe OS input device? EyeOS is the design: MediaPipe iris landmarks become a smoothed gaze vector that moves the Windows cursor through SendInput, and a WebSocket feeds the same stream to a three.js scene.
→ The safety model is the part to keep: the cursor moves only when four gates all pass, an explicit opt-in, an ESC kill switch, a focus check and an ACTIVE state. The prototype never ran end to end, so calibration and latency are unmeasured.
EcoNode
What does a security-first starting point for a crop-genomics screening tool look like? I directed Google’s Antigravity agent through the scaffold: JWT refresh rotation in Redis, argon2 hashing, login rate limiting, security headers and segmented Docker networks.
→ Shelved it the same day, before the model layer got past a stub: a hardened shell around a mocked core.
LinkLeap AI
I wrote the pitch deck for LinkLeap AI, a set-it-and-forget-it job agent for new grads: connect a profile, upload a master résumé, name a target role, and the agent finds openings, tailors the résumé per role, applies and surfaces interviews.
→ That month I built a one-afternoon job-matching prototype. Ten months later I rebuilt the idea properly as JobApp, directing Claude Code, and reversed the shortcuts: public ATS feeds, your own sign-in, a truthfulness gate on every tailored line, review before submit.
Tinkering
I bring up ESP32-S3 and LoRa hardware on my own bench. Three ESP32-S3 boards have connected to it, a LilyGO T-Deck among them, on the Arduino and Espressif toolchains with CP210x and CH340K USB-UART drivers, and the Meshtastic 2.7 release is staged for flashing.
→ For bring-up I wrote my own I²C diagnostic in Arduino C++: it pulses the OLED reset line, sweeps addresses 1–126 every 3 seconds and reports each device that ACKs over serial. A working mesh node is the next milestone.
How close to a printer’s limit can a model go? I scaled a 1.96-million-triangle statue scan to three printers’ build heights and sized a glider so its diagonal wing plate fits each bed, then ported Bambu projects to an Anycubic Kobra 3 V2, TPU included.
→ All six scaled variants land within 4.3 mm of their printer’s limit, and each statue scale is the largest 0.01 step that fits. The models are third-party; my part is the fit, the supports and the slicer settings.
On the first night of CruzHacks 2026 I sketched a voice-agent front end in one 34-minute session: a glowing orb driven by a Web Audio analyser that pulses with microphone volume, wired to the ElevenLabs conversational SDK, with typed input as a fallback.
→ No agent was ever connected. The team’s entry that weekend became SlugBites, which won two category awards.