Research
4 papers and essays, plus the coursework behind them.
Papers & essays
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noise
Typed-decision models answer in fixed-shape fields, such as a yes or no, one option from a list, or a score, in a single forward pass, and they are sold as unable to hallucinate. I tested what that guarantee leaves out. BEMA is an open, from-scratch reproduction of the interface: one small transformer encoder with binary, 77-way and regression heads, trained only on human-labeled data and calibrated with temperature, Platt, isotonic and split-conformal methods. On clean held-out text it looks trustworthy: temperature scaling cuts the binary head’s expected calibration error from 0.0159 to 0.0046. Then I added two adjacent-character typos per query, the ordinary fast-typing kind. Accuracy on the 77-way head fell from 81.9% to 51.2% (n = 3,080), while mean confidence fell only from 0.791 to 0.612. The same gap appears on a second, 151-way dataset. Retraining on typo-augmented data halves the accuracy drop, and the gain carries over to typo types it never trained on, but it does not close the gap. The output guarantee is real: the model cannot emit an answer outside its schema. It can still be confidently wrong on input squarely inside its own domain. A noise-robustness number belongs next to ECE whenever a model’s confidence is sold as a safety signal.
Read
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systems
A language model is most dangerous when it is confidently wrong, and its own confidence is a weak guard. This position paper describes an architecture I built nine times between April and August 2026. Agents do the work, deterministic checks that no model can vote on test it, a council of models from other labs challenges it, and a human takes every split decision. From those builds come three working theories: adjudication beats synthesis, the model that writes should not check its own work, and agreement is not proof. Only the third has data behind it so far. In a two-model cross-check of 94 form-field mappings, all 3 known errors sat inside the 7 disagreements, though the agreements were never audited. On 20 planted citations, a string matcher and an LLM judge agreed on 19 and were both wrong on 2 of those, so their failures overlapped instead of cancelling. Recent work finds that errors correlate across model providers too. I therefore treat independence as a property to measure, not a feature to buy, and propose an evaluation on three tasks with checkable ground truth: error rates when checkers agree and when they disagree, adjudication against synthesis, and self-review against cross-lab review.
Read
Killing Good Ideas
Most of what I start, I stop. This essay lays out the kill rules I actually use and the builds that taught them to me: cutting a feature I could not build honestly, throwing out results that looked too good, keeping the engine when the product goes quiet, and a lesson I learned backwards, that the market map belongs before the code.
Read
Uncorrelated Failure Modes
The design rule behind most of what I build: never trust one source of truth. Independent methods vote, deterministic code does the arithmetic and keeps the books, every fact carries its provenance, and a person decides whenever the checks disagree. This essay explains the rule through my own systems, including where it has already failed me: checks that were supposed to be independent, and weren’t.
Read
Academic work
- Generative AI: How It Works, Why It Fails, and Why It MattersSTS (Science, Technology and Society)
An explainer on how large language models work and why they fail. They are trained to predict the next token, not to check the truth, so fluent output can still be wrong. The paper defines hallucination through NIST AI 600-1’s “confabulation” and walks through documented failures, from fabricated citations to the sanctions in Mata v. Avianca. It covers how training data shapes errors and lays out a four-step routine for verifying model output. Then it turns to my own systems. The multi-agent and council designs I build supply three working theories: adjudication beats synthesis, the writer should not be the checker, and agreement is not proof. BEMA, the model reproduction I architected and ran by directing AI coding agents, supplies a measured result: two adjacent-character typos cut a classifier’s accuracy from 81.9% to 51.2% (n = 3,080) while its mean confidence fell only from 0.79 to 0.61. All eight references are real and checkable.
- U.S. Agricultural Subsidies and the World’s Poorest: Closing for the OppositionECON 3
Closing speaker for the CON side of a team debate on the motion that “American subsidies on agricultural products are harming the poorest people in the world and should be immediately eliminated.” A closer can’t introduce new evidence, only reframe it, so I held the motion to its exact wording. Our evidence showed real damage from U.S. farm support: jobs lost to the sugar program, deadweight loss, and 4 billion dollars in WTO-authorized EU tariffs over the Boeing dispute. That damage lands on U.S. industries and rich trading blocs, not on the world’s poorest. My closing reframed the case around incidence. Most least-developed countries are net food importers, so cheaper world food helps them, and immediate elimination would hit them with a price shock. I also separated crop insurance and disaster aid from the subsidies that drive overproduction.
- Architecting the Social Fabric: Pluralism, Urban RedevelopmentCTW 2
A literature review that synthesizes three sources into one argument: cohesion in a diverse society is built deliberately, through spaces and processes, and does not follow from proximity. Patel’s case for civic “builders” sets the frame. Eck’s distinction between diversity, a demographic fact, and pluralism, active engagement, sharpens it. A 2025 mixed-methods study of the Jordan Downs public-housing redevelopment in Watts, Los Angeles, grounds it: 21 stakeholder interviews and a survey of 647 residents found higher social cohesion there than at two comparison sites. The review closes with implications for planners and policymakers. I then built it into a single-page HTML and CSS research poster sized for a 48 × 36 inch board, with a source-synthesis table, a conceptual framework diagram and a print stylesheet.
- Daily Step Count and Student LifestyleOMIS 40 (Statistics & Data Analysis I)
A survey study with a 7-person team: what predicts how much Santa Clara students walk? We surveyed 29 students on their self-reported daily steps and eight candidate drivers: distance from campus, class load, attendance, transportation, exercise frequency, social outings, job status and year. Then we worked the conditionals. Students who exercise at least three times a week cleared 8,000 steps 60% of the time (9 of 15), and students who go out at least four days a week did so 55% of the time (5 of 9). We closed on the limitations: a sample under 30, a convenience sample, self-report bias, outliers and missing confounders.