How we made a text-to-speech model respond in sub-50 ms

toebee 129 points 32 comments August 21, 2026
nari-labs.com · View on Hacker News

Discussion Highlights (11 comments)

toebee

time-to-first-audio (TTFA) is critical for realtime voice applications. open source implementations (e.g. vLLM-Omni, SGLang-Omni) are often too slow for production and can have issues with realtime playback if you push for lower latency. we wanted to fix that. we optimized qwen3-tts, a popular OSS TTS model, to achieve 34 ms p95 TTFA at 10 requests per second on 1 x H100. we open source the implementation and benchmark, as well as a breakdown of how it was done. github: https://github.com/nari-labs/nari-qwen3-tts

dominotw

chatgpt responds super fast but says filler words like 'hmm..' 'let me think' and responds later with delay

bellowsgulch

GPT‑Realtime‑2 is really weird. Perhaps just because it's bidirectional and now has the failure mode as a possibility, it responds too soon with filler at awkward times, and it's generally overeager. I feel like there was plenty of opportunity to just work on latency engineering like this effort.

nowittyusername

This is right up my alley as ive been building a local voice agent for a year now. Ive tried many different models and have a custom implementation for omni voice that ive tuned for over many months. Ive never been able to achieve faster then 200ms ttfa for that model at 24 steps, but the reason is .... quality. I find that there is a lot of room for improvement in many tts models out there by a huge margin. But there is also a quality hard wall that you eventually hit that the tradeoff of faster latency but lower quality is not worth it. When making a really well sounding voice agent quality of voice, cadence, expression, etc... matters a lot. It will be interesting to try this implementation and see if its quality outputs match my expectations, if so great job indeed.

armcat

Having built my own voice assistant ( https://github.com/acatovic/ova ) and having tried many other services and models, I feel the real win is when this is on-device, and by "on-device" I mean being very inexpensive to run on a phone, and not H100. I've now been using Pocket TTS which is super fast, and also Chatterbox and Fish Audio S2 Pro (on the Mac/PC), I feel we are so close, yet so far. The quality is amazing, but can we take this to the next level and make it run on mobile? What would it take?

zuzululu

this is cool but for agent scenarios unless an LLM bakes in the speech tokens directly, the latency is lost to inference, and this is what makes openai's voice model so interesting also sweet spot is under 150ms so the remainder is inference latency turn around, a 50ms turnaround including tts-stt would ofc be the dream that is "this ai agent is indistinguishably present and sentient" area

MaxikCZ

no video demonstration?

TZubiri

Of course speed is good, but if you don't add an artificial latency (or better, use the extra time for some QA, guardrails, etc...), the model will come off as creepy at best, and the conversation will feel awkward for the user. Humans have a roughly 200ms auditive processing latency, (audio input to neural response), in conversation we know and account for this, such that if someone responds in 100ms, we interpret that we interrupted them and that their message doesn't come in response to what we just said, but what we said before. This can be especially relevant in sentences where an interruption would sharply contrast. "I think murder is bad, but.." If someone cuts of right after the but, a human would interpret that the interjection responds to the fact that someone thinks murder is bad. Which is starkly different than interrupting someone after they are about to excuse murder.

jmesmith

any plans to make this available on cloudflare ai workers (or similar)? Looks super cool, I'd love to try it!

cfferrys

looks good!

vivzkestrel

- a bit unrelated but still had to ask - when recording gaming footage with OBS studio with my mic plugged in, i want to convert my voice to a tts type voice in real time - Basically I speak in my tone but the output is one of your GPT voices - Anyone know of a library or plugin that can accomplish this in real time

Semantic search powered by Rivestack pgvector
4,128 stories · 37,281 chunks indexed