Work / Fintech & RegTech

BlueCollarPal

Procurement and virtual-card control plane for trade contractors: text a request, approve it, pay with a one-time card.

Status
in-build
Role
Solo: directed AI coding agents through two builds from a system-architecture spec (OpenAI Codex, then Claude Code)
Timeline
Jun 2026 – Jul 2026
  • Python 3.12
  • FastAPI
  • Pydantic v2
  • pytest
  • Next.js 15
  • React 19
  • TypeScript
  • Expo / React Native
  • Postgres (row-level security)
  • Lithic single-use cards (adapter, mocked)
  • Twilio webhook validation
  • OpenStreetMap Nominatim and Overpass

TL;DR

  • A procurement control plane for mechanical, electrical and plumbing contractors: a foreman texts a materials request, the office approves it, and the supplier gets paid with a single-use virtual card for exactly that order.
  • The rule it is built on: models extract; deterministic code controls money.
  • Approval is idempotent and role-gated, can mint exactly one card per order, and refuses mock cards in production mode. 25 backend tests cover the guards.

The problem

On a jobsite, materials get ordered by phone call and paid with a shared company card. The office finds out what was bought when the statement arrives, and nobody checks the delivery ticket against the request. Every step where money moves is a manual step, and every manual step is a place for over-ordering, wrong parts and card misuse.

What I built

A foreman texts a request from the jobsite. The system turns it into a bill of materials, prices supplier carts, checks the project budget, and stops for an office manager’s approval. It then issues a single-use virtual card for exactly that order and reconciles the delivery ticket against what was asked for.

I directed AI coding agents through two builds. The first, built with OpenAI Codex over 15 prompts in June 2026, is a FastAPI backend with:

  • a 15-state order state machine;
  • money in integer cents throughout, with a basis-point card buffer;
  • role-gated, idempotent approval that can issue exactly one card per order;
  • a single-use-card adapter keyed for idempotency, and a production mode that refuses to mint mock cards;
  • SMS webhook signature validation, with texts from unregistered numbers rejected and logged;
  • supplier-cart selection that needs at least 70% coverage, then the highest confidence, then the lowest price;
  • reconciliation that flags short, over, unrequested and over-priced lines.

Around it sit a 14-page Next.js dashboard for the office, an Expo mobile app for the field, and a Postgres schema with row-level security on all 18 tables, written for the next stage. Every external provider (card issuing, SMS, supplier pricing, OCR) runs mocked or unconfigured in this build; the money controls around them are real, tested code.

The second build, with Claude Code in July 2026, adds a blueprint-takeoff prototype that sets an estimator agent against a skeptic agent. On a planted panel schedule, the skeptic flagged that the schedule listed 18 twenty-amp positions while its note said 16, caught the estimator’s arithmetic error, and recommended a request for information. A deterministic reconciler, not either agent, sends every challenged line to human review.

Key decisions

  • Decision: models extract, code controls money. Why: a parsing mistake should cost a review, never a payment. Trade-off: every money path needs explicit rules and tests.
  • Decision: one single-use card per approved order. Why: the card can only buy what was approved, once. Trade-off: changes after approval need a new approval and a new card.
  • Decision: an explicit state machine for orders. Why: every transition is named and checked. Trade-off: adding a state means updating the transition table and its tests.

The hard part

Making “approve” safe to press twice. Approval is where a model-built request becomes spendable money, so it has to survive retries and double taps, be gated by role, and never mint a second card.

The code keys approval on an idempotency key, checks for an existing card before issuing, and refuses mock cards in production mode. Downstream, card-provider webhooks are deduplicated by event id and over-limit spends are declined. Each guard has its own test: approval is idempotent and issues one card, over-limit spends are rejected, production refuses mock cards, and duplicate webhooks are ignored.

Results

  • 25 backend tests, with no failures on the last recorded run.
  • A local end-to-end run: register, simulate an SMS request, approve, all returning success.
  • The estimator-versus-skeptic run caught a planted 18-versus-16 discrepancy and an arithmetic error, and routed the line to a human.

What I’d do next

  • Wire the Postgres schema and its row-level security into the runtime, replacing the file-backed store.
  • Connect sandbox card issuing and SMS, then run the flow end to end against real providers.
  • Put the skeptic on a different vendor’s model from the estimator, so the two don’t share blind spots.
  • Code: private, walkthrough on request.
  • The estimator-versus-skeptic pattern: The Council

Verified numbers

MetricValueSource
Backend tests, no failures on the last recorded run25local: BCP/backend/tests/test_workflow.py
States in the order state machine15local: BCP/backend/app/core/status_machine.py
Tables with row-level security (28 policies, written for the next stage)18 of 18local: BCP/database/migrations/
Pages in the office dashboard14local: BCP/frontend/app/**/page.tsx
m.mittal
Home
Work
Research
Lab
About
Experience
Now
Résumé
Uses
Colophon
Contact
AccountWardin-build
Redlbuilt
The Councilresearch
BEMAresearch
EXIT LIQUIDITYin-build
NEXUSarchived
Warrant + Forgebuilt
Concordarchived
BlueCollarPalin-build
Chispenbuilt
Warrant Portalbuilt
JobAppin-build
Night/Dayexplored
BLACKSITE: NULLin-build
SlugBitesbuilt
Riptidebuilt
AeroBitesarchived
Internbuilt
CAD & 3D printingexplored
ForgeCouncilexplored
ADDE: adversarial due-diligence engineexplored
Model routing in practiceexplored
AI due-diligence market mapexplored
Quorumarchived
Colossus Wakeexplored
FitFindrbuilt
Project Omniexplored
Oblivionexplored
MeetWisebuilt
EyeOSexplored
LinkLeap AIexplored
Up-Toexplored
Offline speech-to-notes (Java)explored
AI-assisted game production pipelineexplored
COLLAPSEarchived
THE TRIALSarchived
EcoNodearchived
ESP32-S3 / LoRa benchexplored
CleanPlaybuilt
Rezonyrbuilt
Habit Tracker Telegram Botbuilt
Job Board Aggregatorin-build
Visual Hand Trackbuilt
CLI Task Trackerbuilt
All work44 entries
Can’t Hallucinate, Can Still Be Wrong: Calibration of a Typed-Decision Model Under Input Noisepaper
Disagreement as Signal: A Hybrid Multi-Agent and Council Architecture for Error Detection in LLM Systemspaper
Killing Good Ideasessay
Uncorrelated Failure Modesessay
ADDE: adversarial due-diligence engineexplored
AI due-diligence market mapexplored
AI-assisted game production pipelineexplored
Aura Chatexplored
Colossus Wakeexplored
EcoNodearchived
ESP32-S3 / LoRa benchexplored
EyeOSexplored
FitFindrbuilt
ForgeCouncilexplored
LinkLeap AIexplored
MeetWisebuilt
Model routing in practiceexplored
Offline speech-to-notes (Java)explored
Oblivionexplored
Print benchexplored
Project Omniexplored
Provably-fair outcome enginebuilt
Quorumarchived
Up-Toexplored
Copy hello@mitanshm.com
Switch theme
Play motion
RésuméPDF
Open GitHub ↗mit37
Open LinkedIn ↗
Toggle layout gridh
Toggle single-key shortcuts/ h