Skip to content
Latency (Time-to-First-Token, Sub-Second Voice Response)
voice ai

Latency (Time-to-First-Token, Sub-Second Voice Response)

Latency in conversational AI is the delay between when a user finishes speaking and when the system begins to respond. Time-to-first-token (TTFT) is the model's share of that delay; in voice agents the full round trip must stay near one second to feel natural.

A voice turn has several stages, and each adds delay: detecting that the caller has stopped speaking (end-pointing), transcribing the audio, sending the text and context to the language model, waiting for its first token, generating the rest, converting the text to speech, and streaming that audio back over the phone network. Time-to-first-token measures only the model wait, which is why a vendor quoting an impressive TTFT can still deliver a slow call. The number that matters to the caller is the end-to-end gap from their last word to the first sound of the reply, measured on a real phone line, not in a browser demo.

In human conversation, a pause of more than about a second reads as confusion or a dropped line, and on a phone call in Saudi Arabia or Egypt the caller will say 'ألو؟' and start repeating themselves, which then collides with the agent's late reply. Latency also compounds with Arabic: dialect speech recognition can be slower than English, and long polite openings push more audio through the pipeline. Streaming at every stage (transcribe while the caller speaks, start speaking before the model has finished writing) is what brings the round trip under a second, together with hosting the pipeline close to the telephony region.

Nano AI's Arabic voice agents are built for sub-second responses over existing phone numbers and are monitored after launch, since latency drifts as models, traffic and telephony routes change. When you test any vendor, do not ask for their TTFT figure; call the demo line from a mobile in your city, count the silence after each of your sentences, and ask what they measure in production and what happens when the model provider is slow.

Chat on WhatsApp