Engineering
Voice Agent Interruption Handling: Barge-In in Production
Abhishek Bahukhandi

A candidate cuts our interviewer off mid-question — "wait, can I just use a hash map?" — and the agent keeps talking for another second and a half before it notices. That second and a half is the entire problem. Voice agent interruption handling looks like one if branch in a message handler, and is really three separate queues, two clocks that disagree, and a conversation history that quietly stops matching what the person heard.
Taqari runs live voice interviews on two transports: OpenAI's Realtime API over WebRTC for the behavioural round, and Gemini Live over a raw WebSocket for the coding round. We have already written about why we ship both transports. This post is the layer above that — what barge-in detection actually detects, what you must cancel when it fires, and how to stop the transcript from lying to the model on the very next turn.
Three things break at once when someone barges in
It helps to separate the failure modes, because they have different fixes and different symptoms.
The agent talks over the candidate. Detection fired, the server stopped generating, and the browser is still playing audio it received before the interruption. This is the one users notice and the one teams usually fix first.
The agent answers a question nobody finished asking. Detection fired too eagerly — a cough, a "mm-hmm", a chair scraping — so the turn ended while the candidate was still mid-thought.
The agent references something it never said out loud. The model generated four sentences, the candidate heard one and a half, and nothing told the server where playback actually stopped. Two turns later it says "as I mentioned, the constraint is sorted input" and the candidate has no idea what it is talking about.
The first is a queue problem. The second is a tuning problem. The third is a state problem, and it is the one that is easiest to ship without noticing.
Barge-in detection: the microphone never stops streaming
The precondition for any of this is that you keep sending microphone audio while the agent is speaking. If you gate the upstream on "agent is idle", you have built a walkie-talkie and no amount of downstream cleanup will save it. Our coding round pumps 16 kHz Int16 chunks out of an AudioWorklet continuously, from the moment the session opens until it closes, regardless of what the agent is doing.
Why the detector lives on the server
Both providers run voice activity detection server-side, and that is the right place for it. The server already has the audio, and more importantly it is the only component that knows the state of its own generation — so it can stop producing tokens at the instant it decides the person is speaking. A browser-side detector would be a second opinion arriving over the network, racing a decision the server has already made.
In the Gemini Live protocol the client asks for this explicitly in the setup message. The SDK types describe START_OF_ACTIVITY_INTERRUPTS as: "If true, start of activity will interrupt the model's response (also called 'barge in')." We set it, along with TURN_INCLUDES_ONLY_ACTIVITY so that silence between utterances is not folded into the candidate's turn.
realtimeInputConfig: {
automaticActivityDetection: {
disabled: false,
endOfSpeechSensitivity: 'END_SENSITIVITY_HIGH',
prefixPaddingMs: 150,
silenceDurationMs: 2000,
},
activityHandling: 'START_OF_ACTIVITY_INTERRUPTS',
turnCoverage: 'TURN_INCLUDES_ONLY_ACTIVITY'
}
The behavioural round asks OpenAI for the same behaviour in different words — turn_detection of type semantic_vad with interrupt_response: true and a low eagerness, which trades a little detection latency for far fewer false triggers while someone is thinking out loud.
The dials that decide what counts as an interruption
These three fields do most of the work, and they fail in different directions. Tune them against recordings of real sessions, not against yourself saying "hello, hello" into a laptop in a quiet room.
prefixPaddingMs
How much audio before the detected speech onset gets included in the turn. Too low and you clip the first consonant, so "sort it first" is transcribed as "ort it first" and the model answers something adjacent. We sit at 150 ms. The cost of padding is latency you pay on every single turn, so this is not a dial to be generous with.
silenceDurationMs
How long the candidate has to stop talking before the turn is considered over. This is the most context-dependent value in the whole config. In a coding interview people pause for several seconds mid-sentence while they stare at an editor, and an agent that jumps in after 500 ms of silence is unbearable. We run 2000 ms in the coding round for exactly that reason — it is deliberately patient.
endOfSpeechSensitivity
How aggressively the detector calls the end of an utterance. High sensitivity ends turns sooner, which makes the agent feel responsive and makes it more likely to cut someone off mid-thought. We pair high sensitivity with the long silence window above: quick to recognise that speech stopped, slow to act on it.
The one rule worth remembering
Detection is the cheap half. The expensive half is that your browser is holding audio the person has not heard yet — and a model that believes they have. Every interruption bug we have shipped lived in that gap.
Cancelling in-flight model output across three queues
When the Gemini server decides it has been interrupted, it sets a flag on the server content message. The protocol definition is unusually direct about whose job the cleanup is: interrupted means "a client message has interrupted current model generation. If the client is playing out the content in real time, this is a good signal to stop and empty the current playback queue."
"Stop and empty" is two operations, and there is a third the server handles for you. Here is what is actually in flight at the moment an interruption lands.
Queue one: the buffer currently playing
A AudioBufferSourceNode is single-use. You call start() once and stop() once, and there is no pause — a stopped node is dead and the next chunk needs a fresh node. That makes the cancel itself trivial, but it means you must hold a reference to the node that is currently playing, not just the queue behind it. Wrap the stop() in a try/catch: if the buffer finished on its own microseconds earlier, stopping it again throws, and an exception here would skip the rest of your cleanup.
Queue two: decoded chunks waiting their turn
This is the one that causes the audible talk-over, and it is invisible in local testing because on a fast connection the queue is usually one or two chunks deep. On a congested network the server can be several seconds of speech ahead of the speakers. Emptying the array is one line; remembering to do it is the whole trick.
if (serverContent.interrupted) {
audioQueueRef.current = []; // queue two
if (currentAudioSourceRef.current) {
try { currentAudioSourceRef.current.stop(); } // queue one
catch (e) { /* already ended */ }
currentAudioSourceRef.current = null;
}
isPlayingAudioRef.current = false; // the mutex
return;
}
Queue three: tokens the model is still generating
This one you mostly do not own, and that is a good thing. Because detection runs server-side, generation has already stopped by the time the flag reaches you. The exception is the manual interrupt — a stop button, or a session with automatic detection disabled — where you have to say so explicitly.
Voice agent interruption handling over WebRTC vs a WebSocket
This is where the two transports stop being interchangeable, and it is the strongest practical argument for WebRTC that we know of.
Over WebRTC, return audio is a media track. The browser owns the jitter buffer and the playback clock, so when the server stops sending packets, the sound stops shortly after — queues one and two simply are not yours. OpenAI's own Agents SDK documentation draws the line explicitly: WebRTC handles buffered audio cleanup, while "in WebSocket setups you still need to stop local playback yourself".
Over a WebSocket you hold raw PCM in application memory and schedule it yourself, so every queue above is code you write and code you can get wrong. We think this is worth it in the coding round for other reasons — the transport post covers them — but barge-in is the bill you pay for that choice.
Keeping the transcript consistent after an interrupt
Now the part that is easy to ship broken, because nothing sounds wrong in the session where it happens. It sounds wrong two turns later.
What the model remembers versus what the candidate heard
The model's context contains the full turn it generated. The candidate's memory contains only the audio that finished playing. Everything sitting in queue two when you emptied it is the difference, and if nothing closes that gap the model will confidently refer back to hints it delivered into the void.
The Realtime API exposes the fix directly. Its reference client offers cancelResponse(id, sampleCount), where sampleCount is described as "the number of audio samples that have been heard by the listener", and the call exists to "interrupt the model and prevent it from 'remembering' anything it has generated that is ahead of where the user's state is". The Agents SDK does the same thing a level up: on the WebSocket transport it listens for input_audio_buffer.speech_started and "truncates the assistant audio to what the user actually heard".
Note what that requires of you: a running count of how many samples actually left the speakers. Not how many you received, and not how many you decoded — how many played. If you are queueing audio yourself, that number has to be maintained by your playback loop, and it is the single most useful piece of state in the whole subsystem.
The Gemini Live path gives you less rope here. Detection and generation both live on the server, so its context is truncated at the point where it stopped generating — which is usually close enough, but it does not know how deep your local queue was. On a bad connection those two points are not the same, and the drift is yours to measure.
The half-written code block problem
Our coding round has a second transcript, separate from the model's: we accumulate the output transcription client-side and scan it for fenced code blocks so that code the interviewer dictates lands in the CodeMirror editor automatically. That accumulator is reset when a turn completes.
An interrupted turn does not complete. So if the agent is halfway through dictating a function when the candidate barges in, the opening fence and a few lines of a function body stay in the buffer and become the prefix of the next turn's text. The next closing fence the agent emits — possibly for something entirely unrelated — closes a block that started in a turn nobody finished hearing, and a mangled splice lands in the editor.
The fix is one line, and the lesson generalises well beyond our editor: every piece of per-turn state needs a reset path on the interrupt branch, not just on the completion branch. Go through your handler and list what you clear on turn completion. Anything on that list that you do not also clear on interruption is a latent bug waiting for a candidate who talks fast.
The races that cost us the most time
The ended callback that fires after you stopped the source
Calling stop() fires the node's ended handler. If that handler is the thing that advances your queue — ours calls the queue processor to start the next chunk immediately — then your cancel path triggers the very function you are trying to suppress. It is harmless only because we empty the queue before stopping the node, so the processor wakes up, finds nothing, and goes back to sleep. Reverse those two statements and you will play exactly one more chunk of the cancelled turn. Order matters, and it is not obvious from reading either line alone.
The chunk that arrives after the interrupt
The interruption flag and the last few audio messages of the cancelled turn are in flight at the same time. A chunk that was already on the wire arrives after you have cleaned up, gets appended to a now-empty queue, and plays — one short, orphaned syllable from a sentence that was abandoned.
Clearing the queue cannot fix this, because the problem is arrival order, not queue contents. The robust answer is an epoch: keep a turn counter, increment it on every interruption, stamp each chunk with the counter that was current when it arrived, and drop anything stamped with a stale value in the playback loop. That is the improvement highest on our own list.
What we would do differently
If we were rebuilding voice agent interruption handling from scratch, three things, in the order we would do them.
- Make the playback position first-class. Track samples actually played, in one place, as the authoritative clock for the whole subsystem. Everything else — truncation, metrics, the epoch counter — is easier once that number exists.
- Gate chunks by epoch rather than by queue state. Cheap, removes the arrival-order race entirely, and makes the cleanup path idempotent so a duplicate interruption signal costs nothing.
- Log every interruption with how much audio was discarded. The distribution of that number tells you whether your detection is too eager and whether your queue is too deep, and you cannot tune either dial honestly without it.
None of this is exotic. It is bookkeeping — but it is the bookkeeping that separates an agent that feels like a conversation from one that feels like a voicemail system. If you want to hear where we have got to, the free mock interview is the fastest way to interrupt it yourself, and the Realtime API WebRTC walkthrough covers the session setup that sits underneath all of this.
Frequently asked questions
What is barge-in in a voice agent?
+
Barge-in is when a person starts speaking while the agent is still talking, expecting the agent to stop and listen. Handling it means detecting the speech, cancelling the audio still queued or playing in the browser, and making sure the conversation history reflects only what the person actually heard.
Should barge-in detection run in the browser or on the server?
+
On the server, in almost every case. The server already receives the microphone stream and is the only component that knows the state of the model's generation, so it can stop producing output the moment it detects speech. Browser-side detection adds a second opinion that can disagree with the first.
What do you have to cancel when an interruption fires?
+
Three things, not one. The audio buffer currently playing, the decoded chunks still sitting in your application queue, and the model's own in-flight generation. Miss the second and the agent keeps talking for a second or more after the person interrupted it.
Why does the transcript drift after an interruption?
+
Because the model remembers everything it generated, while the person only heard what finished playing. Unless you tell the server where playback actually stopped, its context contains sentences nobody heard, and it will reference them in the next turn as if they were said.
Does WebRTC handle barge-in automatically?
+
Largely yes. Return audio arrives as a media track the browser owns, so when the server stops sending, playback stops with it. On a WebSocket you hold the decoded chunks yourself, so stopping playback and emptying the queue is application code you have to write.
How sensitive should voice activity detection be?
+
Sensitive enough to catch a real interruption within a few hundred milliseconds, blunt enough to ignore a cough or a backchannel. Tune the padding before detected speech and the silence required to end a turn separately — they trade off against different failure modes.