Work / AI Systems
JobApp
Local-first job-search OS: finds roles, tailors résumés behind a truthfulness gate, fills forms, waits for your OK.
- TypeScript (strict ESM)
- pnpm workspaces
- Claude API (Opus, Sonnet and Haiku tiers)
- Playwright
- PGlite
- Drizzle
- zod
- Gmail and Calendar OAuth
- Deepgram and ElevenLabs
- Stripe REST
- Vitest
- GitHub Actions
TL;DR
- A local-first job-search operating system: it finds roles on the company boards you watch, tailors a résumé for each one behind a deterministic truthfulness gate, fills the application in a browser you’re signed into, and by default stops for your go-ahead before anything is submitted.
- Every rewritten bullet has to cite the master-résumé bullets it came from, and a checker rejects any number, skill, employer, title or date those sources don’t contain.
- I architected, specified and directed Claude Code through the build in under a week: a 13-package TypeScript monorepo with about 22,000 lines of source and 335 tests passing in CI.
The problem
AI résumé tools drift toward inflated numbers and borrowed skills, and “beat the ATS” pressure pushes toward keyword stuffing. Auto-apply tools go further and submit on your behalf. I wanted the speed of automation with the opposite defaults: nothing invented, and nothing sent without a human looking.
What I built
- Discovery. Public job feeds from six applicant-tracking platforms (Greenhouse, Lever, Ashby, SmartRecruiters, Workable and Workday), deduplicated, filtered on preferences and scored against a fit floor.
- Tailor loop. A deterministic selection of the most relevant roles, bullets, projects and skills, then up to three refine iterations. Each iteration must pass the truthfulness gate, then a Claude judge, and the best accepted version wins.
- ATS simulator. Free and deterministic, with no ATS API. It scores every draft from 0 to 100 on weighted must-haves, keywords, title, years and education, and feeds the gaps back into the refine loop.
- Render and re-parse. The best version is rendered to PDF and DOCX, then parsed back to check that a parser reads what a recruiter sees.
- Applying. Playwright in a persistent browser profile, with autonomy as a dial: assist, review-then-submit (the default) or autopilot. Autopilot is refused outright on the most complex enterprise systems, and one-click apply flows are assist-only. Every run saves evidence screenshots. On a CAPTCHA, a bot challenge or a login wall, it hands control back to the user instead of trying to get past it.
- Around the pipeline. A Gmail classifier that files replies and moves applications along, interview prep with voice mock interviews, referral ranking from an exported connections file, and a locked-down local dashboard with single-use sign-in links, signed sessions, CSRF protection and a strict CSP.
- Model routing. Judgment goes to Claude Opus, extraction to Sonnet and high-volume classification to Haiku, with per-model pricing kept for cost accounting.
Key decisions
- Decision: a strict subset rule for tailored bullets. Why: a single invented number is worse than a missed keyword. Trade-off: some legitimate rewordings get rejected; I chose false rejections over false claims.
- Decision: review-then-submit as the default. Why: an application goes out under your name, so a human sees it first. Trade-off: it’s slower than autopilot, on purpose.
- Decision: local-first on embedded Postgres. Why: your résumé, applications and inbox stay on your machine. Trade-off: two processes opening the same data directory silently lose writes, so every opener takes a cross-process lock.
The hard part
Making résumé tailoring provably non-fabricating. JobApp makes provenance structural. A tailored bullet is valid only if it names the master bullets it derives from. The check is extractive, not stylistic: the numbers and hard-skill terms in the new bullet must be a subset of those in its sources, and identity fields (employer, title, dates, education, contact) must match the master exactly.
This runs before any model judgment, so a judge can’t be talked past it. The judge covers what extraction can’t see, which is meaning drifting while the facts stay put. The loop keeps the best version that passed, so a failure falls back to the deterministic selection rather than to a fabricated résumé, and rendering to PDF and DOCX then re-parsing closes the loop on format.
Results
- 335 tests passing in CI on main, with the typecheck green.
- Form filling verified against seven synthetic ATS-style fixture pages, including one CAPTCHA wall.
- Built between September 17 and 22, 2026.
What I’d do next
- Run it end to end on live postings, which is the next milestone.
- Measure the truthfulness gate’s false-rejection rate on real tailoring runs.
Links
- Code: private repo, walkthrough on request.
- The same refuse-rather-than-guess rule: AccountWard · Warrant + Forge
Verified numbers
| Metric | Value | Source |
|---|---|---|
| Tests passing in CI on main, across 21 files | 335 | github.com/mit37/JOBAPP/actions/runs/35770890424 |
| Packages in the monorepo, plus a CLI | 13 | mit37/JOBAPP/packages/ |
| Job-board feed adapters | 6 | mit37/JOBAPP/packages/discovery/src/adapters/ |
| ATS-simulator match weights (must-haves · keywords · nice-to-haves · title · years · education) | 45 · 20 · 10 · 10 · 10 · 5 | mit37/JOBAPP/packages/ats-sim/src/match.ts:60 |
| Synthetic ATS-style pages the form filler is verified against | 7 | mit37/JOBAPP/packages/executor/src/fixtures/ |