πŸŽ™οΈ Project Showcase: OpenVoice Flow

1/11/2026 at 1:31:50 PM • ~5 min read

For this blog post, I want to share a project I have been working on recently called OpenVoice Flow. It is a cross-platform, open-source real-time voice transcription application.

OpenVoice Flow is now open-source! You can find the repository on GitHub.

Why I built it

I have always wanted a fast voice transcription app that keeps the audio on my own machine. I tried Aqua Voice and a few others, but they were either too expensive or sent everything to a server I did not control. So I decided to build my own.

It is built with Tauri 2, which means a Rust backend and a webview UI, with React and TypeScript on the front. Transcription happens as you speak, the API keys are your own, and adding another provider is a small amount of work.

What it does

  • Real-time transcription with low latency
  • Multi-provider support - Soniox (tested), OpenAI Whisper, Ollama (experimental)
  • Cross-platform - Runs on Windows, macOS, and Linux (Tauri powered)
  • Modular architecture - Easy to add new transcription providers
  • BYO-API-key model - Bring your own API keys, no vendor lock-in
  • Privacy-first - Data stays on your device, local key storage
  • Floating overlay - Live transcription feedback in a draggable bubble
  • History view - Access past transcriptions with audio and re-transcription capability
  • Personal dictionary - Custom word replacements, formatting rules, and instructions
  • Network diagnostics - Latency and stats per provider
  • Global hotkeys - Toggle or Push-to-Talk recording modes

What it runs on

  • Frontend: React 18 + TypeScript + Vite
  • Backend: Rust (Tauri 2)
  • Audio: cpal for cross-platform audio capture
  • Database: SQLite for transcript storage
  • Build System: Vite for frontend, Cargo for Rust

What took longest

A few parts took longer than I expected:

  • Cross-platform audio capture - Different OSes handle audio differently. Used cpal to abstract platform differences, but still had to handle device enumeration, sample rate conversion (44.1kHz vs 16kHz), and channel handling (stereo to mono).

  • Real-time streaming - WebSocket connections can be flaky. Implemented a streaming worker pattern with separate threads for audio processing, control channels for start/stop signals, and reconnection logic with exponential backoff.

  • Global hotkeys - Need to work even when the app is in the background. Used tauri-plugin-global-shortcut which handles macOS (Carbon framework), Windows (RegisterHotKey API), and Linux (X11/Wayland protocols).

  • Overlay window management - The floating overlay needs to stay on top and be draggable. Separate Tauri window with alwaysOnTop, skipTaskbar, and decorations: false. Known limitation: it does not work correctly in macOS full-screen mode.

  • Provider abstraction - Different providers have different APIs. Created a TranscriptionProvider trait that all providers implement, making it easy to add new ones.

What I took away

What I took away from it:

  • Tauri 2 is powerful - Type-safe IPC, great plugin ecosystem, small bundle size, and native performance.
  • Rust’s ownership model prevents bugs - No null pointer dereferences, thread safety enforced at compile time, memory leaks are rare.
  • Async in Rust is powerful but complex - Need to understand Send and Sync bounds, channel types, and Arc<Mutex<T>> for shared state.
  • Audio processing is tricky - Sample rates matter, channels need conversion, and both voice activity detection and chunking have to be right before anything downstream works.
  • TypeScript + React is great for UI - Type safety and component reuse, with hooks keeping the state handling simple.
  • Testing is hard - Need to mock audio devices and APIs for integration tests, but manual testing is still required.
  • Privacy is a competitive advantage - Users care about where their data goes. Local storage, no vendor lock-in, optional audio saving, and clear history.

Where it stands

OpenVoice Flow is currently in active development and not yet production-ready.

Known limitations:

  1. macOS full-screen mode: the overlay does not work correctly, which is a macOS window management restriction
  2. Credentials: API keys stored in plain text (encryption planned)
  3. Provider testing: Only Soniox is properly tested
  4. Platform support: Only tested on macOS, Windows/Linux planned

What comes next

What I am planning to work on:

  • Windows and Linux testing and bug fixes
  • Encryption for API key storage
  • More transcription providers (Google, AWS, Azure)
  • Improved VAD for better silence detection
  • Export/import of transcripts
  • Plugin system for custom providers

Wrapping up

Rust and Tauri turned out to be a good fit for this. The performance is there, the bundle is small, and the compiler caught most of the threading mistakes I would otherwise have shipped.

The GitHub repository is MIT licensed, so it can be used for personal or commercial projects.

I hope you enjoyed this blog post and I will see you in the next one!

Get in touch