Skip to content

The thesis

What speech is actually for.

Read a transcript of a conversation you were in and you will barely recognise it.

The words are all there. What is gone is most of what you meant — which parts you were sure about, which ones you were only checking, where you were about to stop and let the other person in. You did not decide any of that. You did it with timing and pitch and breath, and the person listening read it without deciding either.

Say “fine” four different ways and you have said four different things, and only one of them is the word. Everyone hears the difference, and nobody was taught it.

None of that survives being written down. We have taught machines to read very well, and reading is not how people explain themselves to each other.

Our founder, on why we are working on this:

The biggest mistake people make with AI speech is thinking it's about making robots sound pretty. It's not. Speech is the highest-bandwidth protocol humans have for transferring complex intent. If you want a frontier model to actually run the physical world, it can't just process code—it has to master the underlying frequency of interaction. We aren't building a voice tool. We are building the conscious protocol that lets autonomous systems finally understand, adapt, and operate in our reality.

What follows from it

Then the test has to change.

Speech systems are graded on how natural they sound. You play clips to a panel of listeners, ask each of them to rate what they heard, and average the scores. It is a real measurement and it answers a real question, and the question is whether the voice is pleasant.

It never asks whether anything happened. The test that decides whether a machine can work through speech is what the person on the other end does next: whether they answer, whether they read the number back, whether they do the thing they were asked. A voice that sounds beautiful and gets hung up on has failed, and it will still score well.

We do not pass that test everywhere yet. It is the one we build against, which is why the work is not finished when the audio sounds good.

If you are building on this

Write the read, not only the words.

The usable half of this is smaller than the thesis. How a line is delivered is a product decision, the same as what the line says. Where it pauses, which word it leans on, how slowly it reads out a number someone has to write down — those change what the listener does next, and they are yours to set rather than ours to guess.

So set them, and say the whole thing while you are at it. Nothing you send costs more than anything else you send, which is the one part of this we have already finished.

A protocol is only worth what runs on it.

If you are building something that has to be understood out loud, ask for a stream on the production API and tell us what it is for.