This page is written to be looked things up in rather than read end to end. Controls are named exactly as the app names them in English; if you run the interface in another language, the layout is identical and the position of each control is the same.
Two things apply everywhere and are not repeated in each entry. Most settings take effect immediately — the Save button at the bottom of Settings is there for the few that do not, and the line beside it says so. And nothing on this page changes what happens to your audio; that is covered on how it works and privacy.
| If you are looking for… | Section |
|---|---|
| The scenario tiles, languages, Start | The main window |
| Output device, sound test, volume, ducking | Audio and output |
| Subtitles and the on-screen caption window | Screen |
| Anything behind the gear icon | Settings, tab by tab |
| Two-way calls | Meeting-only controls |
| Old sessions, exports, summaries | History and exports |
1. The main window
The window has three parts: a left rail of settings for the scenario you picked, a stage in the middle where the translation appears, and a top bar carrying status and the icons that open everything else.
The two scenarios
The pair of tiles at the top of the left rail is the first choice you make, and it decides what the app captures. Picking one redraws the rail below it.
- Video & Game (“one way”) — translates sound coming towards you and plays the translation in your headphones. Nothing of yours is sent anywhere. This is the scenario for videos, streams, games, podcasts and any call where you only need to understand the other side.
- Meeting (“two way”) — runs two translations at once: what the other side says, translated for you, and what you say, translated and spoken into the call for them. A line under the tiles states the requirement: “Meeting mode needs a virtual microphone to send your translation to the other side.”
Meeting will not start on Windows without a virtual microphone. This is the one place where the in-app wording undersells the rule: the start is blocked, not degraded, and the app raises a “Two-way meeting (optional)” card instead of beginning a half-working session. Install a free virtual cable such as VB-CABLE, restart Voxis, and the scenario becomes available. Meeting also needs a paid plan — see pricing.
Source — System audio or Input device
Source appears on the Video & Game scenario only and has two buttons.
- System audio — the default. Everything this computer plays is translated, with the app's own translated voice excluded so it never ends up translating itself.
- Input device — translates a sound source plugged into the machine instead of the computer's own playback: a games console or set-top box wired into the line input, a television's audio output, a mixer, a USB audio interface, a radio. Choosing it reveals an Input device picker directly below, listing the recording devices Windows knows about.
Two things behave differently on the Input device path, and both are deliberate. The original sound is not lowered while the translated voice speaks — the ducking control does not apply — and there is no echo gate, so use headphones or a separate output so the translation cannot feed back into the input you are capturing.
Windows' own Listen to this device is invisible to System audio. If you are monitoring a line input that way, you must pick Input device; left on System audio the app hears nothing at all. An audio interface running in ASIO mode is also not listed — switch it to its shared Windows driver so it appears.
Meeting ignores the Source choice entirely. It always captures the system audio, whatever is selected on the Video & Game rail.
The language pickers
Both pickers sit above the translation stage.
- I hear in — the language the translation is spoken and captioned in. The hint under it says as much: “The language you'll hear the translation in.” This is the only picker in Video & Game.
- They hear in — Meeting only. The language your own words are translated into for the other participants.
- The arrow between them is the swap button, Meeting only. It exchanges the two languages in one step and restarts the session.
There is no “from” picker. The note under the pickers — “spoken language is detected automatically” — is the whole answer: the engine identifies the language being spoken for you.
Two notes can appear under a picker. “Voice choice isn't available in this language yet” means the target is served by an engine that does not accept a chosen voice. “Not available on the free plan…” and “Free voice: subtitles only in this language” are free-plan notes and do not appear on a paid plan.
Microphone pickers
The app shows a microphone where it is relevant, and all of them are the same setting — change it in one place and the others follow. On the Meeting rail it sits under Your voice and is the microphone you actually speak into. On the Video & Game rail it appears only when Source is set to Input device, and is then the sound source being translated. The master copy lives in Settings › Audio, labelled “Your voice in Meeting; the ‘Input device’ source in Video.” The first entry in every list is (Default), meaning whatever Windows currently considers the default recording device.
Names (Meeting only)
Two short boxes — Your name and The other person's name — sit under Names on the Meeting rail. Anything typed here is handed to the engine as a proper name so it is carried through as written rather than translated or misheard. The hint alongside adds a genuinely useful tip: one-word replies are the easiest thing to mistranslate, so say two words — “Yes, sure” rather than “Sure”.
Start, Stop and the hotkey chip
Start is the single button at the bottom of the stage; it starts the scenario currently selected. The small key combination printed on the button is the global hotkey for that scenario, so you can start it without bringing the window forward. The defaults are Ctrl+Alt+1 for Video & Game, Ctrl+Alt+2 for Meeting, Ctrl+Alt+0 for Stop and Ctrl+Alt+O for the on-screen text; all four can be changed in Settings › Keyboard shortcuts.
Once a session is running, Stop becomes available and two live tools appear beside it.
- Skip — “Skip this sentence”. Cuts the translated sentence currently playing and moves on to the next one. In Meeting it affects only what you hear; it never cuts the voice going out to the other side, which would leave them in unexplained silence.
- Fast mode — captions appear as soon as they exist instead of waiting for the voice to catch up, and if the translated voice falls behind it speeds up to 1.5× without changing pitch. Off by default.
There is no Save button on the stage. The transcript is written when you stop and again every couple of minutes while a session runs.
The status pill and the top bar
The pill in the centre of the top bar is the app's own account of what it is doing, and it carries a level meter with a decibel readout. While idle it reads “Start translation to check the signal”; during a session it moves through connecting, listening and translating, and turns to a warning or error state if something goes wrong. If the meter never moves while sound is clearly playing, the problem is capture rather than translation — the audio setup checklist is the page for that.
Beside the pill, a bar shows the minutes left on your plan. To the right of the top bar sit Report a problem, the theme toggle, Translation history and the gear that opens Settings. Above the translation stage, a Source chip names the input device when one is in use, and a No bot · private badge opens a short explainer of where the audio goes.
Report a problem is worth knowing about before you need it. It sends a small technical snapshot — app version, channel, your Windows version, the audio mode, the quality setting and the language pair — with the app log attached and secrets and your Windows username removed. You can add a category, a severity, a description, and preview the exact payload before sending. Your transcript is included only if you tick that box. You are given a reference code afterwards; quote it if you write to us.
Finally, the Clear stream icon at the top right of the translation stage empties the on-screen conversation. It does not delete anything already saved.
2. Audio and output
The Audio block on the left rail is shared by both scenarios and holds four things.
- Output — shows the device the translated voice plays through, with a Change link that opens Settings › Audio where the full device list lives.
- Sound check — the Open link raises the diagnostics dialog. It is available only when no session is running.
- Translation volume — how loud the translated voice is, from 0% to 150%. 100% is the default; above 100% is a boost for quiet source material.
- Original audio (while speaking) — Video & Game only. How much of the original sound stays audible while the translated voice speaks, from 0% (silent) to 100% (untouched). The default is 30%. This is the control people mean by “ducking”, and it applies to System audio, not to the Input device path.
What the sound check actually tests
The dialog runs three probes, and they are independent of one another — a pass on one says nothing about the others. Run the one that matches the direction that is failing.
- System audio — a bar that moves when the app is hearing what your machine plays. Play a video or some music; if the bar stays flat, the app is not capturing.
- Output device — Play test signal sends a short tone to the output you selected. If you do not hear it, the output device is the problem.
- Microphone — a bar for the microphone you picked. This is the one that matters for the outgoing direction in Meeting.
3. Screen
Two switches sit under Screen on the left rail.
- Subtitles — on by default. Controls whether caption lines appear in the translation stage inside the app window. Turning it off leaves the spoken translation running with nothing written on screen.
- On-screen text — off by default. Opens a separate caption window that floats above your other windows, so you can read the translation over a full-screen video, a game or a call. Its hotkey is Ctrl+Alt+O by default.
The on-screen window shows the translation only — not the source line — and carries the “Powered by Voxis” badge unless you are on a paid plan and have turned it off.
The four looks
The appearance of that window is set in Settings › General, under On-screen text look. A live sample line under the picker shows the result, and changes apply instantly while the window is open.
| Preset | What it looks like |
|---|---|
| Voxis bar | The original Voxis caption bar, reproduced exactly — a rounded dark bar, left-aligned, up to three lines. If you upgraded from an earlier version, this is what you already had and nothing changes until you pick another. |
| Cinema | Large, bold, centred white text on a wide black band near the bottom of the screen, two lines. The film-subtitle look. |
| Minimal | A small, narrow, quiet bar — the least intrusive option, two lines. |
| High contrast | Large yellow text on solid black with a white border, three lines. For readability over busy or bright video. |
Two more controls sit under the preset. Text size offers Small, Medium and Large, and scales whichever preset is selected. Position puts the window at the Bottom or the Top of the screen.
4. Settings, tab by tab
The gear icon in the top bar opens Settings. Seven tabs run down the left of it. The Save button at the bottom commits the few settings that are not applied the moment you change them.
General
Appearance
- Interface language — 23 languages for the app's own buttons and labels. This is not the translation target; it only changes the words in the interface.
- Theme — Dark or Light. The same toggle is on the top bar.
- On-screen text look, Text size and Position — described under Screen above.
- Speaker labels (S1/S2) — on by default. When two different voices alternate, caption turns are tagged S1 and S2, and the tags are written into saved transcripts. It runs entirely on your own machine and adds no cloud calls. Labels only appear once a second voice has actually been heard, so a single-speaker session looks exactly as it did before. The spoken translation still uses one voice.
Advanced
- Write OBS subtitle file — off by default. Writes the current caption line to a plain text file at %APPDATA%\Voxis\obs_subtitle.txt, which you can point an OBS text source at to burn captions into a stream or recording. The file is rewritten only when the line changes.
- Show “Powered by Voxis” badge — on by default, and it is the one visible plan difference in Settings. Hiding the badge requires a paid plan; on any other plan the switch is locked on and hovering it says “Subscription required”. The badge appears on the on-screen text window and in the OBS file.
- Show the tour again — replays the short first-run walkthrough that points at each part of the window.
Audio
- Output device — where the translated voice plays. Headphones are the recommendation for any two-way session.
- Microphone — one setting with two jobs: your voice in Meeting, and the Input device source in Video & Game.
- Sound check — the same three-probe dialog as the rail link.
- Record audio (source + translation, separate files), under Advanced — off by default. Saves each session's source audio and its translated audio as two separate WAV files in your transcript folder, never as a mix.
Dual-track recording is Video & Game only. It switches itself off in Meeting, so the other person's live voice is never written to disk. It is off by default for the same reason: the source track contains a real human voice, which is a larger step than keeping a text transcript. Check what the people you are recording expect, and what the law where you are requires.
Translation
- Profile — Custom, Meeting, Film / Video or Conference. One pick sets the speech-detection tuning and the ducking level together for that kind of material. Changing a related control by hand moves the profile back to Custom.
- Dubbing Voice — off by default. Speaks the translation in a voice close to the original speaker's own, rather than a stock voice. Available on some target languages only.
- Translation voice — Automatic, Female or Male. The voice you hear the video or the other person in.
- My voice — Automatic, Female or Male. The voice the other person hears you in, in Meeting.
- Pro source text, under Source text — paid plans only, and the section is hidden on plans that do not include it. Runs a second recognition line alongside the translation so the source line under the captions, and the saved transcripts, are more complete.
- Terms and proper names — a box for brand, person and product names, one per line, so the engine spells them correctly. Use term=replacement if a term should come out differently. There is an Import from file… button, and a Ready-made term list switch with a Show link that reveals exactly which names are in it.
Two limits on the voice and term controls are worth knowing. Voice gender applies only on the engine that supports a named voice; on the other one the field is ignored, which is why some target languages show “Voice choice isn't available in this language yet” rather than letting the setting do nothing quietly. Dubbing Voice and a chosen gender are mutually exclusive — switching cloning on disables the gender picker, because the cloned voice already carries the speaker's own. And a term list does not reach every language, for the same reason. A term also biases the recogniser towards hearing it, which helps a conversation about those names and hinders one that is not about them — which is why the ready-made technology-brand list is asked about once rather than switched on for everybody.
Saving
- Transcript folder — shows the current path, with Browse…, Open folder and Reset to default. The default is Documents\Voxis\Transcripts, and each session gets its own folder inside it.
- Automatic save formats — TXT, SRT and VTT, all off by default. The JSON record is always saved; these are extra copies written beside it whenever a session is saved, so you do not have to reopen History to export the same formats every time.
Keyboard shortcuts
Four global hotkeys, each shown in a box you can click before pressing the new combination: Video & Game (Ctrl+Alt+1), Meeting (Ctrl+Alt+2), Stop (Ctrl+Alt+0) and On-screen text (Ctrl+Alt+O). They work while other applications have focus, which is the point — you can start and stop without leaving the game or the call.
Membership
Your plan and your balance. The tab shows the active plan and the minutes used, a Need more minutes? card for buying a minute pack, an Invite a friend panel with your referral code, link and statistics, a Manage plan button, a View in Microsoft Store link, and Log out. Members of a team plan see the shared pool instead of a buy button, because a personal pack would not add to it.
About
The version you are running, and links to voxislive.com, pricing, terms, privacy, contact and voice credits. Report a problem is also here, doing the same thing as the chip in the top bar. If an article on this site ever contradicts the app, the version number here is what to quote when you tell us.
5. Meeting-only controls
Everything above applies in Meeting too. These four exist only there.
- The virtual microphone requirement — the outgoing direction needs one, and on Windows the session will not start without it. This is a hard requirement, not a degraded mode. Install a free virtual cable, restart Voxis, then set that cable as the microphone in your meeting app — or pick “Same as system” if your app offers it, in which case Voxis switches it for you. If you select the cable by hand, remember to set your real microphone back afterwards, or your next call without Voxis will have no sound.
- Listen to my translation — off by default. Plays a copy of your outgoing translated voice in your own headphones, so you can hear what the other side is hearing. Use headphones for it: on open speakers your microphone picks the copy back up.
- The two-way transcript — captions and saved transcripts keep the two directions apart, prefixing each turn with Them: or Me: in your interface language, on one chronological timeline. One-way sessions are unaffected and carry no such prefix.
- Swap languages — the arrow between the two pickers, which exchanges “I hear in” and “They hear in” and restarts the session. If the swap cannot be saved, nothing changes.
Meeting also raises a one-off notice the first time you use it, explaining that it both listens to the other party and sends your speech to them, and pointing out that recording or AI-translating a conversation may mean telling the other participants first. You can tick “Don't show again” once you have read it.
6. History and exports
The clock icon in the top bar opens Translation history. Every session is saved to your own disk — the JSON record is written when a session stops and again every couple of minutes while one is running, so an interrupted session is not lost.
- Search history… — filters the list as you type and highlights the matching text. Sessions are listed newest first.
- Starred only — the star beside the search box filters to sessions you have starred. Starring a session also protects it from routine housekeeping, so it is the thing to do with a session you want to keep indefinitely.
- Edit — opens the selected transcript for inline correction; Save writes it back, Cancel discards. Only the text is editable — timings, speaker labels and direction are left alone. Exports read the saved record, so a correction flows through to every format afterwards.
- Delete — removes the selected session and its whole folder, including any exports and audio files in it.
- Open folder — opens the transcript folder in Explorer.
Export formats
Three buttons in the footer: TXT, SRT and VTT. The Bilingual tick beside them decides whether the file carries both the source and the translation or the translation alone; bilingual is the default, and the two variants get different filenames so one never overwrites the other. Exports land inside that session's own folder, beside the JSON. Subtitle cue timings are anchored to the start of the session rather than to the first word, so they line up with a recording of the same session — but remember that simultaneous interpretation runs a few seconds behind the speaker, so the timings are approximate by nature.
Generate summary
Paid plans only. With a session selected, Generate summary produces an AI summary of it; Regenerate replaces one you already have. It is a button you press, never something that happens on its own — pressing it is the step that sends the transcript text anywhere, which is exactly why it is a button. Summaries of Meeting sessions are aware of who said what and are laid out accordingly. If your plan does not include it, the button is not shown.
What is not on this page
- Controls that appear only in the developer build, including the API key section. This page documents the app as distributed from the Microsoft Store.
- The free plan's prompts and walls. They exist, but a paying account does not see them.
- Anything about audio routing beyond which control selects which device — the audio setup checklist is the page for that, and troubleshooting picks up when a control is set correctly and the result still is not right.
If a control exists in your copy of the app and is not described here, that is our gap rather than yours — tell us through contact or the in-app report, and this page gets the correction.