FEATURES

Everything VoxisLive does, in depth.

VoxisLive is a Windows and Linux app that turns any system audio into a natural voice in your language, a few seconds behind the speaker. Here is every major capability, explained.

Driverless system-audio capture

VoxisLive reads your Windows audio mix directly through WASAPI process-loopback — the same low-level Windows Audio Session API that screen recorders use to capture what is playing. There is no VB-CABLE, no virtual sound device, and nothing to route. Install the app and it hears what you hear, immediately, on Windows 10 and 11.

Capture also excludes VoxisLive's own output, so the app never translates its own voice — even in a two-way conversation.

79 languages, swappable mid-session

Pick the language you want to hear from 79 — and in Meeting, the one the other side hears, with a swap button that exchanges the pair in one click without stopping the session. Source-language auto-detection handles multi-language audio.

Two-way meeting mode

Meeting mode runs two live sessions at once: the other party is translated into your language through your speakers, and your own speech is translated into theirs and injected through a virtual microphone. Works alongside Teams, Zoom, Meet, Webex and Discord — and no bot ever appears in the participant list.

VoxisLive Meeting mode translating both directions in the light theme

A native simultaneous interpreter, not a pipeline

Speech goes to a multimodal real-time model that recognizes, translates and re-speaks in a single low-latency pass — the way a human interpreter works in a booth. It begins translating while the speaker is still talking and stays a few seconds behind.

Psychoacoustic ducking

While the translated voice speaks, the original audio is automatically lowered — mirroring professional simultaneous interpretation — then restored when the line ends. You always know who is talking.

Live bilingual transcript & export

Every session produces a searchable transcript — each source line sits directly above its translation. Export it as TXT, SRT or VTT when the session ends.

On-screen subtitles, if you want them

An optional always-on-top caption overlay floats over any app or game, showing the translated line — with S1/S2 tags where speakers change. The bilingual view, each source line above its translation, lives in the app's own caption stream. The spoken voice is the product; captions are there when you need a record.

VoxisLive translating a video with a live bilingual transcript

Private by design

VoxisLive never joins your call as a participant and is not a browser bot. Voice-activity detection, speaker-change labelling and the free tier's local voice all run on your own CPU. On paid plans the captured audio streams continuously to the translation model while a session is live; on the free tier the stream is gated, so only detected speech is sent. No audio is retained after the session. The full capture-to-playback pipeline is published as a source-available excerpt on GitHub, so you can verify exactly what leaves your machine.

Source-available

The desktop engine's audio pipeline is published on GitHub as a source-available excerpt — audit exactly how capture, translation, and playback work. Get the full app from the Microsoft Store with prepaid minutes and zero configuration.

Weighing this against other tools? We keep an up-to-date comparison of the real-time voice translation apps for Windows, including where each one is the better pick.

FAQ

Common questions

01Does VoxisLive need a virtual audio cable?
No — not to capture your PC audio. VoxisLive uses driverless WASAPI process-loopback built into Windows 10 and 11, so there is no VB-CABLE, virtual audio driver or routing utility to install, and your audio setup is left unchanged. Two-way Meeting mode is the exception: on Windows it sends your translated voice through a virtual audio cable such as VB-CABLE and will not start without one; on Linux the app creates its own virtual microphone.
02Is the translation spoken or subtitles?
It is spoken. VoxisLive delivers real-time speech-to-speech translation in a natural voice. A live bilingual transcript and an optional on-screen caption overlay are also available, exportable as TXT, SRT or VTT.
03How far behind the speaker is the translation?
A few seconds, depending on utterance length and network latency. The model starts translating while the speaker is still talking instead of waiting for the sentence to end.
04Which apps does it work with?
Anything that plays audio on Windows: browsers, desktop players, games, and conferencing apps like Teams, Zoom, Meet, Webex or Discord. Capture happens at the OS audio layer, so the source app is irrelevant.
Free to start · 10 free minutes every day

Hear every language, in real time.

Runs on Windows 10 and 11 and on Linux — driverless capture, no setup ritual, no bot in your call.