use case
verified July 2026
The same voice, in a language it never recorded.
Dubbing usually means casting someone else. The voice that built the audience in one market is gone in the next, and what arrives instead is a competent stranger. Cloning inverts that: the identity is a request parameter, so the same speaker says the translated line.
01
Hear the identity survive the translation
Three takes below: one cloned voice, one sentence, three languages. Nothing was re-recorded and no second speaker was cast. The hard case in cloning is a language the reference never spoke, because the identity has to hold with none of the original phonemes to lean on — which is exactly what these three are.
One performance, three languages
Film
One take, one cloned identity, three markets. The dubbing arithmetic — and the same voice in English, Spanish and German — is below.
02
What this is, and what it is not
To be precise about the boundary: Gandr synthesises the dubbed audio and carries the identity across languages. It does not translate your script, separate a voice from a mixed soundtrack, or mux audio back onto video — those stay in your pipeline, where you already have tools you trust.
What changes is the part that used to require a casting session and a studio day per market: the voice itself, available on request, in every language you ship.
- One reference clip, fingerprinted and cached — reuse across a whole catalogue skips re-cloning.
- Per-request emotion and prosody, so a translated line can be delivered at the pacing the scene needs rather than the pacing the words happen to have.
- Unmetered characters, which matters most here: a back catalogue is a large number of characters and a meter turns localisation into a per-title decision.
03
Notes — an engineer's checklist
01Does the cloned voice really hold up in another language?
That is the claim this page is built to let you check rather than take on trust: the three clips above are one cloned reference speaking English, Spanish and German, rendered on the production API. Published, dated specimens for eight languages live at /languages/.
02Do you translate the script?
No. Gandr synthesises speech and carries the voice identity; the translation stays with whatever you already use. Send us the translated text and the reference clip and you get the dubbed audio back.
03Can you separate a voice from a finished mix?
No. There is no source-separation or voice-isolation endpoint on this API, and we would rather say so than let you discover it mid-project. Bring a clean reference clip — five to ten seconds of one speaker is enough.
04What does a back catalogue cost to localise?
The seats you run it on, and nothing per character. That is the whole difference: on a meter, localising a catalogue is priced by its length, so the long tail never gets done.
See also — related sheets
use case
23
TTS for live voice translation
Spex-TTS grew out of a production live-translation product. One cloned identity across 23 languages, 107 ms first audio, measured on the production API.
capability
0
Voice cloning from ten seconds, in the request
Zero-shot voice cloning: a ten-second reference rides inside each request, no training job, and the identity holds across 23 languages.
use case
$0
A narrator who never gets tired of take nine
Script-to-voiceover on a flat seat: re-render the whole video after an edit, in your own cloned voice, without a character count deciding how long the script runs.
capability
7.9
Delivery is a request parameter
The prosody dial moves measured pitch range from 10.1 to 18.0 semitones on one voice, one sentence — continuous per-request control on the production API.
A key and one seat to build this on — the same production API this page measures.
Request access for this use case