L4 · Oct 9, 2026
How we made a voice agent interruptible
- voice-ai
azure-openai-realtime
- latency
- evaluation
- case-study
Our voice agent took about two seconds to react when a caller interrupted it during a knowledge search. Two seconds is a long time to keep talking after someone has asked you to stop.
I lead this realtime contact-center agent at Metrodata. I fixed two separate barge-in bugs, and a controlled A/B on the knowledge-search path measured roughly 0.2 seconds afterward. That test had just n=4 calls per arm. It showed a clear difference in that path, not that we had solved turn-taking.
The later experiments make that distinction important: some changes failed their verdicts, and the live cut-off rate did not improve.
What barge-in means for a speech-to-speech agent
The agent uses Azure OpenAI Realtime. Caller audio streams in and spoken audio streams back. In the default mode, there is no separate speech-to-text or text-to-speech stage.
A Python FastAPI backend bridges the media channel to the model. It also runs tools, including knowledge search over Azure AI Search, and hands calls to a human operator when needed.
Barge-in means the caller starts talking while the agent is speaking, or about to speak. Handling it takes more than cancelling a response:
- Stops the audio the caller is hearing.
- Stops generating the rest of the answer.
- Makes its own memory of the conversation match what the caller actually heard.
The third step is easy to miss. Suppose the model generated a whole paragraph, but the caller heard only half of it. Unless you correct the conversation history, the next answer assumes the caller heard the rest.
Failure one: finished for the model, still playing for the caller
Our cancel logic returned early once the model reported a response as done. That sounds reasonable until you separate generation from playback.
The media channel, Azure Communication Services at the time, could still have audio queued. The model had finished, but the caller was still hearing the agent. An interruption at that point did nothing.
We were also truncating conversation history at the wrong position. We used the number of audio bytes forwarded to the channel, not the amount actually heard. Audio still sitting in a buffer counted as delivered.
I made barge-in playout-aware. A small tracker estimates how many milliseconds of the current response have played to the caller. This is a simplified sketch, not the production code:
on caller speech start:
heard_ms = playout.heard_ms()
channel.stop_audio() # drop queued audio
realtime.cancel() # if still generating
realtime.truncate(item_id, heard_ms)
The order matters: stop playback first, cancel generation if it is still running, then truncate the conversation item at the heard position. The caller gets silence first, and the model's history stops at the same point the caller heard.
I do not have a separate before-and-after measurement for this fix. It shipped as one of the top-priority items from an internal architecture review. The A/B below measures a different code path.
Failure two: the agent could not hear you while it searched
The two-second delay came from the knowledge search tool. It ran inline on the provider receive loop, the coroutine reading every event from the realtime model.
While that loop awaited search results, it could not read the event saying the caller had started speaking. We had made the part responsible for listening wait for I/O.
I moved each tool call into its own cancellable asyncio task, under one 2.5 s time budget. The change also added a single response gate, allowing only one response in flight, and made hand-off to a human operator non-blocking. Again, this is a simplified sketch:
task = asyncio.create_task(
asyncio.wait_for(search(query), timeout=2.5))
# the receive loop keeps reading events
# on barge-in: task.cancel()
Now the receive loop keeps listening during a search, and barge-in cancels the task. I chose one budget rather than nested timeouts so there is one limit to reason about. When it expires, the agent moves on instead of stalling.
How I measured it
I wanted to compare code paths, not two different environments. Our existing voice evaluation harness had a media emulator that fed recorded caller audio into the backend as if it were a real call, a golden question set, and a calibrated LLM judge for answer quality. I added an old-versus-new A/B, using a legacy adapter to run the old path.
Both arms used the same model deployment, search index and caller audio. I ran the primary test from a throwaway runner on Azure Container Instances to keep my laptop's network out of the result. The laptop run remains only as an appendix.
The results on knowledge turns:
- Barge-in latency on the old path: 1,875 to 2,140 ms.
- Barge-in latency on the new path: 93 to 220 ms.
- The new path answered the knowledge question in 4 of 4 calls. The old path escalated to a human in 4 of 4.
With n=4 calls per arm, the roughly tenfold gap repeated, but the sample cannot support percentiles or confidence intervals.
These were harness-driven calls on an experiment runner, not customer traffic. I have no production barge-in measurement to quote. And the test covered knowledge turns only, not ordinary conversation.
What the data did not support
Later, we moved the voice channel to self-hosted LiveKit. A different problem appeared: the agent was talking over callers. The live cut-off rate was 53.3%.
I first suspected that the stop signal reached the media layer too late. I built a local LiveKit reproduction with monotonic-clock traces and tested both a fake model and the real one. Before running it, I wrote down the verdict rules.
The traces mostly refuted my hypothesis: only about 80 ms of audio was heard after the clear. The actual causes were the provider committing the caller's turn during natural pauses, and stale responses that were never cancelled.
I shipped a start-hold and a stale-response cancel, then measured live again. The cut-off rate was 60.0%. It had not gone down.
My notes do not record the number of turns behind either percentage, so these are directional figures. They still do not show an improvement. I recorded a negative result rather than calling the problem fixed.
That experiment series also tested an acknowledgement clip during knowledge search and a rule asking the model to speak a short first sentence earlier. Both failed their pre-registered verdicts. Serving prefetched results could not be tested cleanly, so I reported it as blocked.
A few days earlier, a separate pre-registered check had rejected two proposed changes to latency defaults, each on its own and both together. The defaults stayed unchanged.
None of those rejected or blocked experiments was switched on. I later deleted the levers rejected by the cut-off experiments instead of leaving dormant switches in the codebase.
What I would check first
If you are debugging barge-in, separate what the model generated, what your backend sent, and what the caller heard. Truncate at heard time, and stop playback before dealing with generation or context.
Then check the receive loop. Slow I/O does not belong there. Put tools in cancellable tasks with one time budget, and measure the old and new paths in the same conditions, away from your laptop. Publish the sample size.
Each of these fixes started with a written spec. I directed AI coding agents for much of the implementation, then reviewed and merged their work. But a spec and a merged PR are not proof that a fix helped. Set the verdict rules before testing, accept the negative results, and remove what the data rejected. Our cut-off problem is still not solved.