Experience

The full timeline. Newest first.

  1. –

    Independent research

    BEMA: stress-testing a “can’t-hallucinate” decision model

    An open, MIT-licensed reproduction of a typed-decision model design, built to test the claim that a model with no free text to generate can’t hallucinate. I specified the experiments and directed Claude Code agents through the build and write-up.

    • Two adjacent-character typos cut BANKING77 accuracy from 81.9% to 51.2% (n = 3,080), while mean confidence only fell from 0.79 to 0.61.
    • The gap held on a second domain: CLINC150 went from 71.8% to 45.0% accuracy while confidence fell from 0.75 to 0.56 (n = 5,500).
    • Reported the negative result: on the current run, temperature scaling made the 77-way head’s calibration worse (ECE 0.0210 → 0.0371); isotonic recalibration brought it to 0.0113.
    BANKING77 accuracy under two typos (n = 3,080)
    81.9% → 51.2%
    Mean confidence under the same typos
    0.79 → 0.61
    CLINC150 under typos (n = 5,500): accuracy · confidence
    71.8% → 45.0% · 0.75 → 0.56
    77-way head ECE: raw → temperature → isotonic
    0.0210 → 0.0371 → 0.0113
  2. Pitched Redl to investors

    Redl, my local-first AI workbench

    Pitched Redl to investors in San Francisco (Jul 2026).

    Pitched to investors in San Francisco

  3. Intern, team build with Fluxxion88

    Loop Engineering Hackathon, AWS Builder Loft

    Competed on a team with my collaborator (Fluxxion88). We built Intern, which trains an agent like a new hire: from one hand-made example, an LLM loop writes, runs, scores and repairs a script until it matches, then hands back a script with no model inside.

    • The team’s recorded run climbed from a 68% to a 96% to a 100% match in three attempts.
    Recorded run, attempts 1 → 2 → 3
    68% → 96% → 100% match
  4. Warrant + Forge, two-person team with Fluxxion88

    Estate-settlement AI hackathon

    Built Warrant + Forge with my collaborator (Fluxxion88) over two days, with Warrant forked from Concord’s verifier, ingestion and provider layers: an estate-settlement engine where no fact reaches the ledger without a verbatim quote verified in its source document. It didn’t place, but I opened the AccountWard repo the day after.

    • I directed AI coding agents through Warrant, the engine half; Forge, the form-filling half, is my collaborator’s.
    • Warrant caught 41 of 41 rewritten source quotes before they reached the ledger.
    Rewritten source quotes caught before the ledger
    41 / 41
  5. –

    Software Engineering Intern

    early-stage fintech startup (stealth)

    Worked across the stack on data workflows and user-facing features, directly with the founders on architecture and product decisions.

  6. NEXUS, solo entry, Wargaming track

    SCSP AI+ Expo Hackathon (Phase 1, San Francisco)

    Entered the Wargaming track solo with NEXUS, an AI wargame adjudication engine for supply-chain resilience analysis, and demoed it live to the track’s national-security judging panel, who probed the simulation logic and the AI adversary.

    • Solo entry: I architected and specified NEXUS and directed AI coding agents (Cursor and Antigravity) through the build: a Monte Carlo cascade simulation argued over by a chain of AI agents, streamed to a 3D globe, plus a companion Android app.

    Demoed live to the Wargaming track’s national-security judging panel

  7. –

    Riptide: live AirSim drone mode and volunteer Android app

    Hack for Humanity 2026, Santa Clara University

    My team built Riptide, a platform for coordinating ocean-pollution cleanup. I architected and directed an AI coding agent through a live Microsoft AirSim mode behind its drone console, and through a volunteer companion app for Android on the final morning.

    • A Python Flask bridge flies a simulated quadcopter over RPC: telemetry, camera, depth and LiDAR feeds, and threaded waypoint missions.
    • The simulator’s own log records three bridge-driven missions.
    • Both pieces ran in my build; the team’s submission kept its map-based simulation.
    Bridge-driven missions in the simulator log
    3
    Bridge endpoints
    14
  8. SlugBites: 2 awards

    CruzHacks 2026, UC Santa Cruz

    Our team’s SlugBites won two single-winner category awards: Most Ambitious Project and MLH Best Use of Solana. It’s drone-delivery ordering for UC Santa Cruz dining halls, built on live-scraped menus, with a Python flight layer written for Microsoft AirSim.

    • I wrote 20 of the repo’s 23 commits, including the scraper that turns a human-facing nutrition site into a menu API, meal-period pricing and a SOL-priced checkout, plus the AirSim flight layer.
    • Round two of an idea: three months earlier it was AeroBites at the AWS x INRIX Hack.

    Most Ambitious ProjectMLH Best Use of Solana

    Commits (mine / total)
    20 / 23
    Submissions at CruzHacks 2026
    88
  9. AeroBites: drone food delivery, v1

    AWS x INRIX Hack 2025, SCU ACM

    Built AeroBites on a five-person team: a React ordering app for campus drone delivery, with an AWS serverless back end written as infrastructure-as-code (Lambda, DynamoDB, S3 and a Bedrock function). I was the top committer, including every commit that changed code.

    • We had no drone, so on the final morning I stood up Microsoft AirSim’s Unreal Engine environments and flew a quadrotor in simulation.
    • Three months later the idea came back as SlugBites, on live data, and our team won two awards with it at CruzHacks 2026.
    Team size
    5
    Commits (mine / total)
    13 / 18
  10. –

    Triple major: Finance (BS), Computer Science (BA), Economics (BA)

    Santa Clara University, Leavey School of Business

    Class of 2029. Most finance people can’t ship software, and most engineers don’t follow where the money moves. I’m studying all three sides: finance, computer science and economics.

    • Statistics and conditional probability (OMIS 40)
    • Trade and subsidy economics (ECON 3)
    • Capital-structure strategy in a shareholder-value simulation (BUSN 70)
    • Information systems and web development (OMIS 34)
    • Fall 2026: writing on how generative AI works and why it fails, with my own AI systems as the case studies (STS)
    • Competed at SCU ACM’s AWS x INRIX Hack (Oct 2025) and SCU’s Hack for Humanity (Feb 2026)
    Majors
    3
    Class of
    2029
  11. –

    Vice Branch & Fundraising Lead

    DevFinTech

    Ran fundraising and corporate-partnership outreach for financial-literacy programs.

m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h