Work / Collaborations
Intern
An LLM loop that turns one hand-made example into a deterministic script with no model inside.
- Python
- pandas
- FastAPI
- Pydantic v2
- SQLite
- Server-Sent Events
- pytest
- Next.js 15
- React 19
- TypeScript
- Tailwind CSS 4
- Docker Compose
- OpenAI, Anthropic and Bedrock SDKs
TL;DR
- A collaboration with Fluxxion88. At the Loop Engineering Hackathon at AWS Builder Loft in San Francisco (July 17, 2026), our team built Intern: train an agent the way you’d train a new hire.
- A non-technical user describes a recurring CSV chore and uploads one output they made by hand. An LLM loop writes a pandas script, runs it in a sandbox, diffs the result cell by cell against the example, and repairs it until it matches.
- What you get is a plain script with no model inside, and the model never grades itself.
The problem
Recurring spreadsheet chores are too small to hand to an engineer and too fiddly to trust to a chatbot every week. Asking a model to do the chore each time means paying for it each time and trusting it each time. The better artifact is a script that does the chore deterministically, and the hard part is getting a model to write one that is actually right.
What was built
The user describes the chore, answers three questions, approves a read-back spec and uploads one example output. Then the loop runs:
- Generate. The model writes a pandas script. It never sees the target file.
- Run. The script runs in a subprocess with an empty environment, its own process group, CPU, memory and file-size limits, and a wall-clock kill.
- Score. A deterministic scorer maps columns in four passes, aligns rows on a discovered key, detects row order with a longest-increasing-subsequence, and emits typed findings: unit, rounding, date format, totals row, missing rows and more.
- Repair. The repairer fixes the highest-weight finding first, and the loop stops at a perfect match, a plateau or the budget.
- Guard. An AST check confirms that every attempt, and the finished tool, makes no network or model calls. The UI shows the verdict as a “No AI inside” card.
Around the engine: a FastAPI service with SQLite and a replayable server-sent-event stream, a Next.js 15 UI with a live training ledger and a public file-drop page for each trained tool, and Docker Compose. The build was spec-driven: a 2,906-line, nine-document spec written for AI coding agents, then 19 commits in 2 hours 12 minutes on event day, all pushed by my collaborator and all co-authored by Claude.
Key decisions
- Decision: the model never sees the target and never grades itself. Why: a loop graded against a file it can read will learn to print the answer. Trade-off: every repair has to go through typed findings, not free-form feedback.
- Decision: typed findings instead of one score. Why: “the dates are in the wrong format” is something a repairer can act on; “72%” isn’t. Trade-off: the scorer is tuned to tabular, CSV-shaped chores.
- Decision: ship a script, not an agent. Why: the chore runs forever at zero model cost, and anyone can read what it does. Trade-off: a new kind of chore needs a new training run.
The hard part
Keeping a self-improving loop honest. A code-generation loop graded against a target file is easy to game: the model can print the expected answer as literals. Intern closes that off in four layers. The model never sees the target. The grader is deterministic code that explains each failure as a typed finding. A hard-coding detector treats any target value that can’t be traced to the inputs or the approved spec as distinctive, and rejects attempts that contain two or more of them; anything the human wrote into the spec is whitelisted, so legitimate constants aren’t punished. Last, a static guard checks that the shipped artifact contains no model or network calls at all.
Results
- The team’s recorded run climbed from 68% to 96% to a 100% match in three attempts.
- 29 test functions across the scorer, runner, guard, repairer, API and end-to-end artifact.
What’s next
- Extend the scorer beyond tabular chores.
- Run trained tools on fresh, unseen inputs and report how often they still match.
Links
- Code: github.com/Fluxxion88/intern-hackathon
- Devpost: Intern
Verified numbers
| Metric | Value | Source |
|---|---|---|
| Match against the hand-made example across three attempts (the team’s recorded run) | 68% → 96% → 100% | Fluxxion88/intern-hackathon/README.md:21 |
| Test functions | 29 | Fluxxion88/intern-hackathon/tests/ |
| Build spec written for AI coding agents, in 9 documents | 2,906 lines | Fluxxion88/intern-hackathon/docs/ |
| Submissions at the event (193 participants) | 65 | loop-engineering-hackathon.devpost.com/ |