Relay-1 · real-time speech translation

Translation fast enough
to hold a conversation.

Relay-1 is our own speech-to-speech model. Audio in, translated audio out, fast enough that two people can interrupt each other and the conversation still works. It runs behind the Mac app and behind the API.

Loading translator…

Speak into it. This is the live model, not a recording.

We publish the benchmark, including where we lose.

Five metrics across seven systems on kyutai/Audio-NTREX-4L, a long-form set covering French, Spanish, Portuguese, and German into English. No system wins on every axis — that's the expected shape for a problem that trades speed against correctness. Relay-1 is top-tier on the axes that decide whether a live exchange works: second on naturalness, within 0.01 on translation quality, and competitive on intelligibility.

MetricGPT-RTGeminiRelay-1
Translation quality
XCOMET-XL ↑
0.730.780.77
Naturalness
UTMOSv2 ↑
2.893.643.27
Intelligibility
ASR-WER ↓
0.110.030.06
First audio
TTFA p50 ms ↓
228932693201
Speaker similarity
WavLM ↑
0.780.700.46

Relay-1 doesn't clone the source voice, which is why its speaker-similarity score sits low. For a live conversation we've optimised for clear, natural, accurate output over reproducing each speaker's timbre — a different tradeoff than our dubbing model makes.

Read the full benchmark, all seven systems
Licensing & deployment

Some teams can't route audio through a hyperscaler.

We own Relay-1 end to end, which means we can have the conversations the platform vendors won't: on-premise deployment, data residency in a region you choose, white-label licensing, and tuning against your own domain audio.

Google's Live Translate API lists at $0.023 per minute and is still in public preview. Ours is $0.02, generally available, and you can run it inside your own perimeter. If you need to own the stack rather than rent it, we should talk.

Talk to us about licensing