Work / Collaborations

Intern

An LLM loop that turns one hand-made example into a deterministic script with no model inside.

Status
built
Role
Team member
Collaboration
Built with Fluxxion88
Timeline
Jul 2026 – Jul 2026
  • Python
  • pandas
  • FastAPI
  • Pydantic v2
  • SQLite
  • Server-Sent Events
  • pytest
  • Next.js 15
  • React 19
  • TypeScript
  • Tailwind CSS 4
  • Docker Compose
  • OpenAI, Anthropic and Bedrock SDKs

TL;DR

  • A collaboration with Fluxxion88. At the Loop Engineering Hackathon at AWS Builder Loft in San Francisco (July 17, 2026), our team built Intern: train an agent the way you’d train a new hire.
  • A non-technical user describes a recurring CSV chore and uploads one output they made by hand. An LLM loop writes a pandas script, runs it in a sandbox, diffs the result cell by cell against the example, and repairs it until it matches.
  • What you get is a plain script with no model inside, and the model never grades itself.

The problem

Recurring spreadsheet chores are too small to hand to an engineer and too fiddly to trust to a chatbot every week. Asking a model to do the chore each time means paying for it each time and trusting it each time. The better artifact is a script that does the chore deterministically, and the hard part is getting a model to write one that is actually right.

What was built

The user describes the chore, answers three questions, approves a read-back spec and uploads one example output. Then the loop runs:

  • Generate. The model writes a pandas script. It never sees the target file.
  • Run. The script runs in a subprocess with an empty environment, its own process group, CPU, memory and file-size limits, and a wall-clock kill.
  • Score. A deterministic scorer maps columns in four passes, aligns rows on a discovered key, detects row order with a longest-increasing-subsequence, and emits typed findings: unit, rounding, date format, totals row, missing rows and more.
  • Repair. The repairer fixes the highest-weight finding first, and the loop stops at a perfect match, a plateau or the budget.
  • Guard. An AST check confirms that every attempt, and the finished tool, makes no network or model calls. The UI shows the verdict as a “No AI inside” card.

Around the engine: a FastAPI service with SQLite and a replayable server-sent-event stream, a Next.js 15 UI with a live training ledger and a public file-drop page for each trained tool, and Docker Compose. The build was spec-driven: a 2,906-line, nine-document spec written for AI coding agents, then 19 commits in 2 hours 12 minutes on event day, all pushed by my collaborator and all co-authored by Claude.

Key decisions

  • Decision: the model never sees the target and never grades itself. Why: a loop graded against a file it can read will learn to print the answer. Trade-off: every repair has to go through typed findings, not free-form feedback.
  • Decision: typed findings instead of one score. Why: “the dates are in the wrong format” is something a repairer can act on; “72%” isn’t. Trade-off: the scorer is tuned to tabular, CSV-shaped chores.
  • Decision: ship a script, not an agent. Why: the chore runs forever at zero model cost, and anyone can read what it does. Trade-off: a new kind of chore needs a new training run.

The hard part

Keeping a self-improving loop honest. A code-generation loop graded against a target file is easy to game: the model can print the expected answer as literals. Intern closes that off in four layers. The model never sees the target. The grader is deterministic code that explains each failure as a typed finding. A hard-coding detector treats any target value that can’t be traced to the inputs or the approved spec as distinctive, and rejects attempts that contain two or more of them; anything the human wrote into the spec is whitelisted, so legitimate constants aren’t punished. Last, a static guard checks that the shipped artifact contains no model or network calls at all.

Results

  • The team’s recorded run climbed from 68% to 96% to a 100% match in three attempts.
  • 29 test functions across the scorer, runner, guard, repairer, API and end-to-end artifact.

What’s next

  • Extend the scorer beyond tabular chores.
  • Run trained tools on fresh, unseen inputs and report how often they still match.

Verified numbers

MetricValueSource
Match against the hand-made example across three attempts (the team’s recorded run)68% → 96% → 100%Fluxxion88/intern-hackathon/README.md:21
Test functions29Fluxxion88/intern-hackathon/tests/
Build spec written for AI coding agents, in 9 documents2,906 linesFluxxion88/intern-hackathon/docs/
Submissions at the event (193 participants)65loop-engineering-hackathon.devpost.com/
m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h