Skip to content

use case

verified July 2026

Every phrase, every speed, every dialect.

A language product is not a script, it is a bank. Every vocabulary word, every example sentence, every dialogue, read at a normal pace and again more slowly, in every language you teach — the better product is the one with more of it. A per-character meter prices exactly the coverage that makes the product good.

01

The bank is the product

A phrasebank multiplies. A few thousand words become tens of thousands of short utterances once you add example sentences, then dialogues, then a slower reading of each, then the same set again in the next language. Under per-character pricing every one of those axes multiplies the same character count again, so the meter charges most for breadth — the one thing a learner actually wants more of.

On an unmetered seat the multiplication stops being an invoice. Adding the slow reading, or the ninth language, or the twelve extra dialogues a level needs costs what the first one did, which is nothing per character. The rational move becomes covering the long tail rather than shipping the top three languages and leaving the rest as text.

02

What a phrasebank needs from the engine

  • One cloned tutor voice across 23 languages, so the same teacher carries every edition — and the accent is whichever native speaker you clone from.
  • cfg_weight loosens the pacing, so a slower reading is a second render of the same line rather than a second recording session.
  • A pronunciation dictionary and spell tags fix loanwords, place names and letter-by-letter drills once per request, not after a learner reports them.
  • Reproducible seeds, so correcting one phrase re-renders that phrase without shifting the thousand around it.

03

Notes — an engineer's checklist

01Can one tutor voice teach every language we cover?

Yes. One reference clip carries the identity across 23 languages, so the localised editions keep the same teacher rather than swapping voices per market. Eight of the 23 carry published, dated specimens at /languages/.

02How do we offer a slower reading without recording it twice?

cfg_weight loosens the pacing per request, so a relaxed pass is a second render of the same line. Both passes are unmetered, and a shared seed keeps the voice matched between them.

03Can we cover a specific regional accent?

The accent rides with the reference clip. Clone from a native speaker of the dialect you want to teach and the identity carries that accent; cloning requires a consenting speaker, and every clip we synthesise also carries an inaudible provenance watermark applied at generation.

See also — related sheets

A key and one seat to build this on — the same production API this page measures.

Request access for this use case