Work / AI Systems
Project Omni
The native Rust half of a hold-to-talk AI dictation app: global hotkey, streamed audio, screen capture, typed-back text.
- Rust
- Tauri 2
- Tokio
- tokio-tungstenite
- cpal
- rdev
- xcap
- enigo
- React 19
- Tailwind
TL;DR
- The idea: hold Cmd/Ctrl+Shift+Space, speak, and the text streams into whatever app has focus.
- I directed Cursor’s agent through the native Tauri 2 / Rust client in one four-phase session, type-checking after every phase.
- The cloud transcription relay was never built, so Omni stays a client-side spike.
What I built
Three threads do the work. A global rdev listener catches the chord. A session loop grabs one RAM-only screen frame (the hook for screen context; nothing touches disk) and opens the microphone with cpal. A Tokio relay streams raw PCM over a WebSocket and uses enigo to type returned tokens into the focused app while the chord is held.
A React 19 settings screen runs native permission probes: microphone, Accessibility and Screen Recording on macOS, and the microphone consent store and UI Automation on Windows. By design the client holds no API keys.
The hard part
Never block the audio callback. Capture runs on a real-time thread, and the network doesn’t.
A bounded 128-slot queue that drops audio when full → the capture callback can never stall behind a slow socket → under a long network stall you lose audio instead of freezing the app. The relay reconnects with backoff from 500 ms to 30 s, so a flaky connection recovers without hammering the server.
Status
Explored. Every phase type-checked, but the transcription relay the client talks to was never built, so there was nothing to run it against end to end.
Verified numbers
| Metric | Value | Source |
|---|---|---|
| Rust lines across 13 files | ~1,032 | local: project-omni/src-tauri/src/ |
| Relay reconnect backoff | 500 ms → 30 s | local: project-omni/src-tauri/src/relay.rs:101 |
| PCM queue, drop-on-full | 128 slots | local: project-omni/src-tauri/src/relay.rs:61 |