use case
verified July 2026
Ship the voice in every market, not the top three.
Localisation budgets are allocated by expected return, so the top three markets get audio and the rest get text. That is a reasonable decision under a meter and an unnecessary one without it.
01
Why the long tail stays silent
Under per-character pricing, every additional language multiplies the same character count again. The markets that would each add a little revenue are exactly the ones that cannot justify the multiplication, so they ship without voice and stay smaller, which confirms the decision that created it.
Here the languages are not the cost — the seats are. Adding the ninth language costs what adding the second did.
One voice, twenty-three languages
Film
Under a meter each added language multiplies the character count again. Here the languages are not the cost.
02
One identity, every edition
- One cloned reference carries one identity across 23 languages, so the brand voice is the same person in every market.
- Eight of the 23 carry published, dated specimens at /languages/ — real output with generation dates, not a claim.
- Per-request emotion and prosody, so a localised line is delivered rather than merely pronounced.
- Unmetered characters, which is what makes the long tail affordable at all.
03
Notes — an engineer's checklist
01Which languages are supported?
Twenty-three from a single cloned reference. Eight carry published, dated specimens at /languages/ so you can hear them before committing; the rest run live on request.
02Does the brand voice survive translation?
That is the specific hard case cloning is measured on — a language the reference never spoke, with none of the original phonemes to lean on. The dubbing sheet publishes one voice in three languages so you can judge it yourself.
See also — related sheets
use case
23
The same voice, in a language it never recorded
One ten-second reference carries one identity across 23 languages, so a dubbed track keeps the original speaker instead of replacing them with a stranger.
use case
23
TTS for live voice translation
Spex-TTS grew out of a production live-translation product. One cloned identity across 23 languages, 107 ms first audio, measured on the production API.
capability
0
Voice cloning from ten seconds, in the request
Zero-shot voice cloning: a ten-second reference rides inside each request, no training job, and the identity holds across 23 languages.
capability
7.9
Delivery is a request parameter
The prosody dial moves measured pitch range from 10.1 to 18.0 semitones on one voice, one sentence — continuous per-request control on the production API.
A key and one seat to build this on — the same production API this page measures.
Request access for this use case