Research / Essay
Killing Good Ideas
Abstract
Most of what I start, I stop. This essay lays out the kill rules I actually use and the builds that taught them to me: cutting a feature I could not build honestly, throwing out results that looked too good, keeping the engine when the product goes quiet, and a lesson I learned backwards, that the market map belongs before the code.
Most of what I start, I stop. I don’t count that as waste. It is the filter working, and the filter is the part of building I trust least to instinct. So I have rules for it. None of them came from a book. Each one came from a specific build, and in each case the kill made the work that survived better.
Cut what you can’t do honestly
At an estate-settlement hackathon, where I built Warrant with Fluxxion88, the plan included a model council: two or three models proposing facts independently, with conflicts sent to a person. It was also second on the list of things to cut if time ran short, with a written condition: it needed two different model providers, because with one, cross-model challenge would be theatre, and it was better to cut it than fake it. It was cut.
What shipped in its place was a deterministic verifier. Every fact needed a verbatim quote that a matcher could find in a document on file. In the benchmark, that verifier caught 24 of 24 altered numbers in quoted text. It caught none of 41 cases where a genuine quote was attached to the wrong value, and that zero went into the published results. A council on stage would have looked more impressive, and it would have been a claim I couldn’t back.
BEMA, my stress test of a “can’t-hallucinate” model design, ran on the same rule. Two experiments could not be run properly in the environment: a pretrained encoder was unreachable, and no language-model API key was available for the speed and cost comparison. Both are reported as not run, not approximated. When a dataset’s licensing turned out to mix several sources with different terms, I dropped it and replaced the whole task rather than working around it. A second candidate was rejected for ambiguous licensing, so that part of the study stopped at one new dataset instead of two to four. Half a result I can defend beats a whole one I can’t.
Kill the result that flatters you
The first version of NEXUS, a wargame-adjudication engine I architected solo over a hackathon weekend and built by directing AI coding agents, said every scenario ended in catastrophe. In the first smoke test, 0.958 of simulated cascades failed. It read like a finding. It was a bug: the edge weights exceeded the node thresholds, so almost any single failure took down the whole chain. I flagged it because the verdict never changed, whatever the input. The model was rebuilt with hand-set profiles and wider randomness in the thresholds, and after that the outputs discriminated between targets. One exported run gave 10.51%. A simulation that always says “catastrophe” is broken, not alarming.
A cross-vendor due-diligence prototype taught the same lesson from the other side. Its first live run, with agent swarms from four providers, exited cleanly after 343 seconds. Exit code zero. Underneath, each department had logged 35 to 63 provider errors, only 4 of its 8 departments produced a valuation, and the financial desk missed a planted $9 million receivable trap. A clean exit is not a result. My later councils added deterministic verification in front of the models.
Keep the parts that survive
Ideas die faster than the code inside them, and the code is usually the part worth keeping.
On one day in July, I forked Redl, my local-first AI workbench, into Quorum, a private diligence app running open-weight models on the user’s own machine. 35 minutes after Quorum’s last commit, I forked it again into Concord to run frontier models from several labs. Quorum’s positioning lasted less than an hour. Its engine didn’t die with it.
Concord is the bigger case. After building it, I directed a multi-agent market teardown to find out whether it already existed. The answer put Concord somewhere between fifth and tenth in a crowded, well-funded field, and found only two pieces defensible: a citation tie-out that blocks unsupported findings, and a per-finding challenge by a model from a different vendor. The tie-out had already gone into Warrant, which I had forked from Concord’s verifier and ingestion layers the day before. Thirteen days later I built the benchmark that could prove or break it: 23 planted citations, faithful and fabricated, all 23 graded correctly. Concord the product has been dormant since that benchmark, on 10 August 2026. The parts are still working.
Map the market before the code
This is the rule I learned backwards. Concord’s teardown came two days after its first commit. AccountWard’s 25-entrant competitive landscape, which I had research agents build, came almost four weeks into its build. Both were useful. Both answered a question I should have asked before writing a line of code: who already does this, and what is left that is worth doing?
So the order is fixed now. The market map comes first, then the code. Anything the map says is commodity gets cut before it costs a week.
What a kill buys
Every cut above bought something specific. Cutting the council bought a verifier with a published boundary. Refusing to approximate bought a paper whose numbers survive a hostile reader. Distrusting a clean exit code bought councils that check the models with rules. Letting a product go quiet kept an engine that is still running in the next one.
The idea dies. The evidence and the parts stay.