HOW IT WORKS

From your speakers to your language in a few seconds.

VoxisLive captures system audio directly through WASAPI loopback, detects speech on-device, translates it with a real-time AI model, and speaks the result back — while automatically lowering the original audio.

Step 1 — Audio capture, without drivers

VoxisLive uses WASAPI loopback — the same low-level Windows Audio Session API that screen recorders use to capture “what's playing.” It is native Windows capability: no virtual audio cables, no driver installation, no audio routing changes. Capture is zero-latency relative to playback and adds no audible artifacts.

Competing tools typically route audio through virtual devices like VB-CABLE, which need driver installs (often admin rights and a reboot) and can conflict with exclusive-mode audio, ASIO drivers or anti-cheat systems. VoxisLive skips that class of problem entirely.

Step 2 — On-device speech detection

The app runs on-device voice-activity detection (VAD) to separate speech from silence, background noise and music — locally, on your CPU, with no network round-trip. Only segments identified as human speech proceed to translation, which cuts latency and protects your minute balance. VAD also tracks VoxisLive's own spoken output so it never re-translates its own voice.

Step 3 — One-pass simultaneous translation

Speech segments go to a multimodal real-time model that handles recognition, translation and voice synthesis in a single low-latency pass — collapsing the three sequential network calls of a traditional pipeline (speech-to-text → translation → text-to-speech) into one. Like a human interpreter in a booth, it begins translating while the speaker is still talking.

Step 4 — Spoken playback with ducking

The translated voice plays through your output device while two things happen in parallel: psychoacoustic ducking lowers the original audio while the translation speaks (mirroring professional simultaneous interpretation), and latency synchronization keeps each translation aligned with its speech segment so long sessions never drift.

TECHNICAL REFERENCES & CITATIONS

Architecture & Standards Documentation

VoxisLive's engine is built on open standards and enterprise AI specifications. Read the underlying documentation:

VoxisLive main window translating a video, showing the live bilingual transcript

Watch it live — switching between five languages in one session

This screen recording shows a real conference talk translated in real time, with the target language switched live, mid-talk, across Spanish, Italian, Turkish, Chinese and Korean — same source audio, five different spoken outputs, no restart between switches. It's the exact WASAPI loopback pipeline described above: no virtual audio cable, no subtitles, the translation spoken back in a natural voice while the original stays audible underneath.

How fast is it, really?

VoxisLive is near-simultaneous, not zero-delay: typically a few seconds behind the original speech, depending on utterance length and network latency. For reference, professional human simultaneous interpreters work two to four seconds behind the speaker — VoxisLive operates in that range or faster. Very short fragments are batched to avoid poor single-word translations.

Meetings without a bot

VoxisLive never joins a call as a participant, never requests host permissions, and never touches the meeting app. It reads the audio already playing to your speakers, which makes it invisible to other participants and identical across Zoom, Teams, Google Meet, Webex and Discord. In two-way mode it also translates your own speech into the meeting language through a virtual microphone.

What leaves your machine

With the managed Store app, audio is proxied to the model and no audio is retained after the session ends. On-device VAD separates speech from silence, background noise and music; on paid plans the stream runs continuously while a session is live, and on the free tier only detected speech is sent.

GETTING STARTED

Two modes, set up in under a minute

Video & Game mode — translate anything playing on your PC

Use this for a movie, stream, YouTube video, or game with spoken dialogue.

  1. Open VoxisLive and pick Video & Game from the left panel.
  2. Choose the language you want to hear under I hear in. You never set the spoken language — it is detected automatically.
  3. Press Start. Nothing else to configure — VoxisLive listens to whatever is already playing through your speakers and starts speaking the translation within a couple of seconds.

There is nothing to install and no separate audio device to pick: VoxisLive captures your PC's system audio directly (see "Audio capture, without drivers" above).

Meeting mode — be heard and understood on Teams, Zoom or Google Meet

Meeting mode is two-way: you hear the other person's speech translated into your language, and — if you want them to hear you translated too — your voice is translated into theirs.

Both pickers are target languages, and neither one is the language being spoken. “I hear in” sets what reaches your ears; “They hear in” sets what reaches theirs. The spoken language is detected automatically, sentence by sentence, in both directions — so whatever language you switch into, it still reaches them in the language you picked, and even if the other side is multilingual or three languages are spoken in the same call, you keep hearing only yours.

  1. Open VoxisLive and pick Meeting from the left panel. Set I hear in (the language you want to hear) and They hear in (the language the other side should hear). Neither one asks what language is spoken — that is detected automatically.
  2. Press Start. You'll immediately hear the other participant translated — this half needs no extra setup.
  3. To be understood back, one extra step is required in your meeting app, not in VoxisLive: open Teams' (or Zoom's, or Meet's) own audio settings and set the microphone to the virtual cable installed on your PC (named "CABLE Output" or similar) — VoxisLive does not install one, it uses whichever is already there. VoxisLive switches your Windows default microphone automatically, but some meeting apps remember whichever microphone you had selected before and don't pick up the change on their own — so double-check it once, at the start of the call.

Skip the last step and Meeting mode still works — you'll hear everyone else translated, you just won't be translated back to them (VoxisLive calls this "listen-only" and says so on screen). No virtual cable installed at all also degrades gracefully to listen-only, never an error.

On a WhatsApp call specifically? See the illustrated WhatsApp call setup guide for exact click-by-click screenshots of both steps.

FAQ

Common questions

01Do I need VB-CABLE or a virtual audio driver?
No — not to capture your PC audio. VoxisLive uses WASAPI loopback, a native Windows API available on Windows 10 and 11, so there is nothing to install or route and no new device appears in your audio settings. A virtual microphone is only optional, for sending your translated voice back into a two-way meeting; without one, Meeting mode runs listen-only.
02Does VoxisLive join my meeting as a bot?
Never. It captures your own system audio locally, so no third attendee appears in Zoom, Teams or Meet, no permission prompt fires, and no platform integration is needed.
03How much delay should I expect?
A few seconds behind the original speech after speech detection and AI processing — the same range as professional human simultaneous interpreters, or faster.
04Does translation quality depend on utterance length?
Yes — longer utterances translate better. Very short fragments are batched or deferred to avoid poor single-word translations.
05Can I choose which AI model translates my speech?
No manual selection is needed or offered. VoxisLive automatically routes each session to the best-performing engine for your chosen language — currently a mix of Google's Gemini Live and Alibaba's Qwen Realtime, described above — and switches to a backup engine automatically if one becomes unavailable mid-session. The interface does not display which engine handled a given session.
06Does it work in a multilingual meeting, whatever language the other side speaks?
Yes. Both pickers in Meeting mode are target languages, not source languages: one sets the language you hear, the other the language the other side hears. The spoken language is detected automatically, sentence by sentence, in both directions — you never choose it. So whatever language you switch into, it reaches them in the language you picked; and even if the other side is multilingual, or three languages are spoken in the same call, you keep hearing the one language you chose.
Free to start · 10 free minutes every day

Hear every language, in real time.

Runs on Windows 10 and 11 and on Linux — no drivers, no setup ritual, no bot in your call.