Research / Essay

Killing Good Ideas

Abstract

Most of what I start, I stop. This essay lays out the kill rules I actually use and the builds that taught them to me: cutting a feature I could not build honestly, throwing out results that looked too good, keeping the engine when the product goes quiet, and a lesson I learned backwards, that the market map belongs before the code.

Most of what I start, I stop. I don’t count that as waste. It is the filter working, and the filter is the part of building I trust least to instinct. So I have rules for it. None of them came from a book. Each one came from a specific build, and in each case the kill made the work that survived better.

Cut what you can’t do honestly

At an estate-settlement hackathon, where I built Warrant with Fluxxion88, the plan included a model council: two or three models proposing facts independently, with conflicts sent to a person. It was also second on the list of things to cut if time ran short, with a written condition: it needed two different model providers, because with one, cross-model challenge would be theatre, and it was better to cut it than fake it. It was cut.

What shipped in its place was a deterministic verifier. Every fact needed a verbatim quote that a matcher could find in a document on file. In the benchmark, that verifier caught 24 of 24 altered numbers in quoted text. It caught none of 41 cases where a genuine quote was attached to the wrong value, and that zero went into the published results. A council on stage would have looked more impressive, and it would have been a claim I couldn’t back.

BEMA, my stress test of a “can’t-hallucinate” model design, ran on the same rule. Two experiments could not be run properly in the environment: a pretrained encoder was unreachable, and no language-model API key was available for the speed and cost comparison. Both are reported as not run, not approximated. When a dataset’s licensing turned out to mix several sources with different terms, I dropped it and replaced the whole task rather than working around it. A second candidate was rejected for ambiguous licensing, so that part of the study stopped at one new dataset instead of two to four. Half a result I can defend beats a whole one I can’t.

Kill the result that flatters you

The first version of NEXUS, a wargame-adjudication engine I architected solo over a hackathon weekend and built by directing AI coding agents, said every scenario ended in catastrophe. In the first smoke test, 0.958 of simulated cascades failed. It read like a finding. It was a bug: the edge weights exceeded the node thresholds, so almost any single failure took down the whole chain. I flagged it because the verdict never changed, whatever the input. The model was rebuilt with hand-set profiles and wider randomness in the thresholds, and after that the outputs discriminated between targets. One exported run gave 10.51%. A simulation that always says “catastrophe” is broken, not alarming.

A cross-vendor due-diligence prototype taught the same lesson from the other side. Its first live run, with agent swarms from four providers, exited cleanly after 343 seconds. Exit code zero. Underneath, each department had logged 35 to 63 provider errors, only 4 of its 8 departments produced a valuation, and the financial desk missed a planted $9 million receivable trap. A clean exit is not a result. My later councils added deterministic verification in front of the models.

Keep the parts that survive

Ideas die faster than the code inside them, and the code is usually the part worth keeping.

On one day in July, I forked Redl, my local-first AI workbench, into Quorum, a private diligence app running open-weight models on the user’s own machine. 35 minutes after Quorum’s last commit, I forked it again into Concord to run frontier models from several labs. Quorum’s positioning lasted less than an hour. Its engine didn’t die with it.

Concord is the bigger case. After building it, I directed a multi-agent market teardown to find out whether it already existed. The answer put Concord somewhere between fifth and tenth in a crowded, well-funded field, and found only two pieces defensible: a citation tie-out that blocks unsupported findings, and a per-finding challenge by a model from a different vendor. The tie-out had already gone into Warrant, which I had forked from Concord’s verifier and ingestion layers the day before. Thirteen days later I built the benchmark that could prove or break it: 23 planted citations, faithful and fabricated, all 23 graded correctly. Concord the product has been dormant since that benchmark, on 10 August 2026. The parts are still working.

Map the market before the code

This is the rule I learned backwards. Concord’s teardown came two days after its first commit. AccountWard’s 25-entrant competitive landscape, which I had research agents build, came almost four weeks into its build. Both were useful. Both answered a question I should have asked before writing a line of code: who already does this, and what is left that is worth doing?

So the order is fixed now. The market map comes first, then the code. Anything the map says is commodity gets cut before it costs a week.

What a kill buys

Every cut above bought something specific. Cutting the council bought a verifier with a published boundary. Refusing to approximate bought a paper whose numbers survive a hostile reader. Distrusting a clean exit code bought councils that check the models with rules. Letting a product go quiet kept an engine that is still running in the next one.

The idea dies. The evidence and the parts stay.

m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h