Skip to the content

IT&SOFT3 min readAI

Voice AI agents in 2026: what already works and what is still marketing

Latency, interruption, recognition in noise and the phone line: where voice agents genuinely apply, and where the demo is prettier than reality.

Voice demos look better than voice rollouts. Not because anyone is lying: a demo is recorded in a quiet room, with a good microphone, on a rehearsed scenario. Reality is a noisy warehouse, an 8 kHz phone line and a person who interrupts halfway through a sentence.

What genuinely works

Taking enquiries on a routine script. Book an appointment, take an order, confirm an address. Short turns, a limited vocabulary, a clear reason for the call. Here a voice agent works, and it is noticeably cheaper than an operator.

Outbound confirmations. Remind someone of a booking, confirm a delivery, ask about a convenient time. The script is entirely in the agent’s hands and there are few possible replies.

Qualifying inbound calls. Work out why the person is calling and pass them to the right specialist with the context already collected. Even if the agent only gathers data and transfers, that saves minutes on every call.

Voice input where hands are busy. A warehouse, a workshop, a car. Often this is not a conversation at all: someone dictates into a form. That is the scenario that works best.

What is still held up by marketing

The long consultation. Five minutes of conversation with branches, objections and references back to what was said at the start. The agent loses the thread, the person loses patience.

Being indistinguishable from a human. Claimed often. In practice people work out that they are talking to a robot within two or three turns—from the rhythm and from intonation that is too even. And that is fine: introducing yourself as a robot is more honest, and the conversation is none the worse for it.

Emotional work. Calming an angry customer, keeping one who is leaving. Here the agent does harm: the feeling of “they would not even give me a human” makes the conflict worse.

Difficult names and addresses by ear. Surnames, street names, contract numbers. Recognition errors in this class of words remain high, and the cost of an error in a delivery address is real money.

Three technical limits that decide everything

Latency. A comfortable pause in conversation is up to 700 milliseconds. Speech recognition, the model and speech synthesis all have to fit into it. Every link adds time, and on a poor network the budget is used up instantly. The practical consequence: the shorter the agent’s turns, the more alive the dialogue feels.

Interruption. People start speaking before you have finished. The agent has to stop immediately and understand what was said. This is implemented separately from the main logic and is one of the main differences between a working system and a demo.

The phone channel. Narrow band, compression, noise. A model that works excellently on a laptop recording makes noticeably more mistakes on a phone line. Testing belongs on a real line. A laptop recording shows nothing.

How we roll this out

We start with one scenario and one number. Not “the company’s voice assistant” but “appointment booking on this number”. We record the conversations and listen to them in full for the first week—not a sample, all of them. We count the share of calls that reach their goal and the share transferred to a human.

The escalation rule is set from the very beginning and is simple: two failed attempts to understand, transfer to an operator. Not five, not “let us try again in different words”. Two.

And separately: the agent always introduces itself as a robot in its first turn. That removes half the irritation and, from what we have seen, does not lower the share of calls that reach their goal.

When to come back to the topic

If you have more than a hundred similar calls a day, the script fits into a minute and the cost of an error is low, now is a good time. If calls come in tens, or every conversation is unique, wait a bit longer—you will lose nothing.

ShareTelegramVK
Author

IT&SOFT

A small team of engineers. We write about the work we do by hand, and about what breaks while we do it. If you have something similar on your plate, write to us and we will go through your case.

Discuss your task
Subscribe to new breakdowns
Next

Got a similar
task?

Describe it in the brief. In working hours we come back with an estimate of time and cost within two hours.