who it’s for
verified July 2026
Gandr for localisation and dubbing studios.
A studio is paid to make the same content work somewhere else. The part that has always cost the most is the voice: a casting session and a studio day per market, repeated per title, which is why back catalogues stay half-localised.
01
Where the length stops mattering
Under per-character pricing a catalogue is priced by its runtime, so the long tail of older or smaller titles never clears the bar. On seats, the languages are not the cost — the parallel capacity is.
The identity travels with the work. One reference clip carries the same speaker into a language the reference never spoke, which is the specific case worth checking rather than trusting.
02
The boundary, stated plainly
Gandr synthesises the audio and carries the voice identity across languages. It does not translate the script, separate a voice from a mixed soundtrack, or mux audio back onto video — those stay in the pipeline you already run.
That is a deliberate line. The engine is a speech API, and a studio has better tools than we do for everything either side of it.
03
Notes — an engineer's checklist
01Does the voice survive a language it never recorded?
That is the hard case, and the dubbing sheet publishes one cloned voice in English, Spanish and German so you can judge it rather than take the claim.
02Can you isolate a voice from a finished mix?
No. There is no source-separation endpoint on this API. Bring a clean reference — five to ten seconds of one speaker in a quiet room is enough.
03How is a catalogue priced?
By the seats you run it through, not by its length. Characters and minutes are unmetered, which is what makes the smaller titles worth doing.
See also — related sheets
use case
23
The same voice, in a language it never recorded
One ten-second reference carries one identity across 23 languages, so a dubbed track keeps the original speaker instead of replacing them with a stranger.
use case
1
Ship the voice in every market, not the top three
Localised product audio on a flat seat: one cloned identity across 23 languages, with the character count no longer deciding which markets get voice.
capability
0
Voice cloning from ten seconds, in the request
Zero-shot voice cloning: a ten-second reference rides inside each request, no training job, and the identity holds across 23 languages.
use case
23
TTS for live voice translation
Spex-TTS grew out of a production live-translation product. One cloned identity across 23 languages, 107 ms first audio, measured on the production API.
A key and one seat to size it on, before anyone talks about a fleet.
Request access for your stack