ποΈ Project Showcase: OpenVoice Flow
For this blog post, I want to share a project I have been working on recently called OpenVoice Flow. It is a cross-platform, open-source real-time voice transcription application.
OpenVoice Flow is now open-source! You can find the repository on GitHub.
Why I built it
I have always wanted a fast voice transcription app that keeps the audio on my own machine. I tried Aqua Voice and a few others, but they were either too expensive or sent everything to a server I did not control. So I decided to build my own.
It is built with Tauri 2, which means a Rust backend and a webview UI, with React and TypeScript on the front. Transcription happens as you speak, the API keys are your own, and adding another provider is a small amount of work.
What it does
- Real-time transcription with low latency
- Multi-provider support - Soniox (tested), OpenAI Whisper, Ollama (experimental)
- Cross-platform - Runs on Windows, macOS, and Linux (Tauri powered)
- Modular architecture - Easy to add new transcription providers
- BYO-API-key model - Bring your own API keys, no vendor lock-in
- Privacy-first - Data stays on your device, local key storage
- Floating overlay - Live transcription feedback in a draggable bubble
- History view - Access past transcriptions with audio and re-transcription capability
- Personal dictionary - Custom word replacements, formatting rules, and instructions
- Network diagnostics - Latency and stats per provider
- Global hotkeys - Toggle or Push-to-Talk recording modes
What it runs on
- Frontend: React 18 + TypeScript + Vite
- Backend: Rust (Tauri 2)
- Audio:
cpalfor cross-platform audio capture - Database: SQLite for transcript storage
- Build System: Vite for frontend, Cargo for Rust
What took longest
A few parts took longer than I expected:
-
Cross-platform audio capture - Different OSes handle audio differently. Used
cpalto abstract platform differences, but still had to handle device enumeration, sample rate conversion (44.1kHz vs 16kHz), and channel handling (stereo to mono). -
Real-time streaming - WebSocket connections can be flaky. Implemented a streaming worker pattern with separate threads for audio processing, control channels for start/stop signals, and reconnection logic with exponential backoff.
-
Global hotkeys - Need to work even when the app is in the background. Used
tauri-plugin-global-shortcutwhich handles macOS (Carbon framework), Windows (RegisterHotKey API), and Linux (X11/Wayland protocols). -
Overlay window management - The floating overlay needs to stay on top and be draggable. Separate Tauri window with
alwaysOnTop,skipTaskbar, anddecorations: false. Known limitation: it does not work correctly in macOS full-screen mode. -
Provider abstraction - Different providers have different APIs. Created a
TranscriptionProvidertrait that all providers implement, making it easy to add new ones.
What I took away
What I took away from it:
- Tauri 2 is powerful - Type-safe IPC, great plugin ecosystem, small bundle size, and native performance.
- Rustβs ownership model prevents bugs - No null pointer dereferences, thread safety enforced at compile time, memory leaks are rare.
- Async in Rust is powerful but complex - Need to understand
SendandSyncbounds, channel types, andArc<Mutex<T>>for shared state. - Audio processing is tricky - Sample rates matter, channels need conversion, and both voice activity detection and chunking have to be right before anything downstream works.
- TypeScript + React is great for UI - Type safety and component reuse, with hooks keeping the state handling simple.
- Testing is hard - Need to mock audio devices and APIs for integration tests, but manual testing is still required.
- Privacy is a competitive advantage - Users care about where their data goes. Local storage, no vendor lock-in, optional audio saving, and clear history.
Where it stands
OpenVoice Flow is currently in active development and not yet production-ready.
Known limitations:
- macOS full-screen mode: the overlay does not work correctly, which is a macOS window management restriction
- Credentials: API keys stored in plain text (encryption planned)
- Provider testing: Only Soniox is properly tested
- Platform support: Only tested on macOS, Windows/Linux planned
What comes next
What I am planning to work on:
- Windows and Linux testing and bug fixes
- Encryption for API key storage
- More transcription providers (Google, AWS, Azure)
- Improved VAD for better silence detection
- Export/import of transcripts
- Plugin system for custom providers
Wrapping up
Rust and Tauri turned out to be a good fit for this. The performance is there, the bundle is small, and the compiler caught most of the threading mistakes I would otherwise have shipped.
The GitHub repository is MIT licensed, so it can be used for personal or commercial projects.
I hope you enjoyed this blog post and I will see you in the next one!