What original-voice dubbing does
Ordinary live translation replaces the speaker. A character shouts in a cut-scene, an actor drops their voice to a whisper, a streamer laughs mid-sentence — and what reaches your headphones is a neutral synthetic voice reading the meaning back to you. The words arrive; the performance does not. For a film or a game, that gap is most of what you came for.
Original-voice dubbing closes part of it. In Video/Game mode, on the incoming leg, VoxisLive can clone the voice of the person currently speaking and render the translation in that voice. The character keeps their own voice while speaking your language: the timbre and register you are hearing on screen, carrying words you understand.
The clone is one-shot. The speaker's voice is sampled once and that sample is reused for the rest of the session, rather than being re-derived for every sentence. Everything else about the session is unchanged — the same capture, the same simultaneous translation running a few seconds behind the source, the same captions and transcript.
Where it works
Three conditions have to hold at once. Miss any one of them and the session runs normally with a stock voice — nothing fails, you simply do not get the clone.
- Video/Game mode. This is a one-way, incoming-only feature. It is not offered in Meeting mode.
- The incoming leg. It applies to the audio your computer is already playing — the film, the stream, the game, the recording. It is never applied to your microphone.
- A voiced target routed to the Qwen engine. VoxisLive routes each target language to one of two engines. Cloning exists on the Qwen path; targets routed to the other engine get the ordinary translated voice.
| Mode | Leg | Original-voice dubbing |
|---|---|---|
| Video / Game | Incoming — what you hear | Available on voiced targets routed to Qwen |
| Video / Game | Incoming — target routed to the other engine | Not available; ordinary translated voice |
| Meeting | Incoming — the other party to you | Not available (Video/Game mode only) |
| Meeting | Outgoing — you to the other party | Not available; your voice is not cloned |
What it is not
The thing people most often expect from the name is the thing that is missing, so it is worth being blunt: your own voice is not cloned on the outgoing leg. In a meeting, what the other side hears is a translated voice speaking your words in their language — not your voice. If you were hoping to join a call and have colleagues hear you, in your own voice, in Spanish, that is not what this does today.
It is also not a post-production tool. VoxisLive works on live audio your machine is playing right now; there is no file or URL to hand it, and no finished dubbed track comes out the other end. If you want a rendered dub of a video you own, this is the wrong tool — see video translation for what the live path actually gives you.
- Not your voice on a call — the outgoing leg is never cloned.
- Not available in Meeting mode, in either direction.
- Not available on every target language.
- Not combinable with the voice-gender choice (see below).
- Not a file converter — live audio only.
The control beside it: voice gender
Next to dubbing sits the other way of choosing how the translation sounds: a voice-gender setting with three values — auto, female, male — set per leg in Settings. It carries the same engine boundary as cloning. The gender choice is effective only on targets routed to the Qwen engine; targets routed to Gemini have no equivalent setting, because that model does not accept a voice selection at all. On those languages the app says the choice is not available, rather than pretending to apply it.
That is a real limitation and we would rather you read it here than discover it mid-film: on a Gemini-routed target, the voice you hear is whatever the model produces. There is no switch, hidden or otherwise, that changes it.
Why you cannot have both at once
A voice clone and a named voice or gender choice cannot be used at the same time. This is not a matter of interface taste — the provider rejects a request that carries both, so a session configured that way would not start at all. The interface therefore makes them mutually exclusive: turning dubbing on takes the gender choice out of play, and picking a gender turns dubbing off.
The reasoning is easy enough once you see it. A clone already carries the speaker's own voice, register included. Asking for a cloned voice and then asking for that voice to be male or female is asking two contradictory questions in one request. So you decide which matters more for what you are watching: the actor's own voice, or a consistent voice you have chosen.
Rule of thumb. Fiction, gameplay and anything where the performance carries meaning: use dubbing. A lecture, a long interview or anything you will listen to for an hour: a fixed, gender-selected voice is usually steadier, because it does not change when the person on screen does.
Which target languages get it
Routing is decided per target language, and it is decided on our side rather than by a setting you flip. VoxisLive translates into 79 target languages; the subset routed to the Qwen engine is where cloning and the gender choice exist. The full list of targets, and what each one speaks with on which plan, is on the languages page.
The practical answer is in the app itself: pick your target, open Settings, and look at the incoming voice controls. If the gender choice is offered, you are on the Qwen path and dubbing is available too. If the app tells you the choice is not available for that language, the target is routed to the other engine and both controls are out of play.
Turning it on
- Install VoxisLive and open it — see download.
- Pick the Video / Game scenario, not Meeting.
- Choose the language you want to hear. This is a target language; the spoken one is detected for you.
- Open Settings and enable the voice clone for the incoming translation.
- Notice that the incoming voice-gender choice becomes unavailable while dubbing is on. That is the mutual exclusion, working as intended.
- Press Start, then play the film, stream or game as you normally would.
If the setting is not offered for your target, the language is routed to the other engine. Changing the target language is the one thing that changes routing from your side. Everything else — capture, pacing, transcripts — behaves exactly as described in how it works.
Honest expectations
Dubbing changes which voice reads the translation to you. It does not change the nature of simultaneous interpretation: the translation still arrives a few seconds behind the source, because the engine has to hear enough of a sentence before it can say anything about it. A cloned voice a few seconds late is still a few seconds late.
It also does not turn a live session into a studio product. The clone path carries one cloned voice through a session, so a scene with several actors does not become several separate cloned voices. If your goal is a finished multi-voice dub, this feature is not it — and we would rather say so than have you find out after buying. What it is good at is narrow and real: one person is talking at you in a language you do not have, and you would like to hear them, not a stand-in.
Plans and minutes are on pricing; the wider feature list is on features.