Engineering
WebRTC vs WebSocket for Voice Agents: We Ship Both
Abhishek Bahukhandi

Taqari runs two kinds of live voice interview, and they sit on two different transports. The behavioural and system-design round talks to OpenAI's Realtime API over WebRTC. The LeetCode-style coding round opens a WebSocket straight to Gemini Live and streams raw PCM up and down. That was not indecision — each provider steers you somewhere different, and running both taught us more than picking one would have.
So this is WebRTC vs WebSocket for voice agents written from inside a product that ships both: what each transport hands you for free, what you end up writing by hand, and which one holds up when the network stops cooperating.
Two rounds, two transports
The shapes are genuinely different. On the WebRTC side we mint a short-lived client secret on our backend, build an RTCPeerConnection, attach the microphone track, and POST the SDP offer once over HTTPS to get an answer back. There is no long-lived signalling channel: one HTTP exchange, then media flows peer-to-endpoint. Control messages — session configuration, transcripts, the events that tell us a turn finished — ride a data channel alongside the audio.
On the WebSocket side, a single call to our backend returns an ephemeral token, and the browser then opens a wss:// connection directly to the provider. The first frame is a setup message carrying the interviewer prompt, the voice, the VAD configuration and the tool declarations. After that, everything — our microphone audio, the model's audio, tool calls, tool results — travels as JSON on that one socket.
WebRTC vs WebSocket for voice agents: the latency budget
The honest headline is that on a clean network, both transports are about one network hop and the model's thinking time dominates everything. Where they diverge is in what sits between the microphone and the wire — and in what happens when the network is not clean.
What the media stack gives you free
When you hand a track to addTrack, you get the browser's whole audio pipeline: Opus encoding at frame sizes tuned for speech, RTP packetisation, an adaptive jitter buffer, packet-loss concealment, and congestion control that lowers bitrate instead of falling over. None of it appears in your source file. The W3C WebRTC specification is long precisely because all of that is specified for you.
The consequence shows up under loss. UDP drops a late frame and moves on; the concealment algorithm papers over a 20 ms gap and the candidate hears a tiny artefact. TCP cannot do that. It must deliver bytes in order, so a single retransmit holds back everything queued behind it, and a WebSocket audio stream on a lossy connection does not degrade gracefully — it goes quiet, then arrives in a burst. That burst is worse than the loss, because now the agent is a second behind the conversation.
What you hand-build on a socket
Our WebSocket round runs an AudioWorkletNode whose processor buffers float samples, converts to 16-bit integers, and posts a chunk to the main thread. Three parameters in that loop are latency decisions.
Chunk size
We settled on 64 ms — 1,024 samples at 16 kHz. That number is a floor on how late the model can learn you started talking, because nothing leaves the browser until a chunk is full. Halving it halves that floor and doubles the frame count, and every frame costs a JSON parse on the far side. WebRTC makes this choice for you at roughly 20 ms, which is a good default nobody has to argue about.
Serialization
PCM is binary and the channel is carrying JSON, so each chunk gets base64-encoded. That is a fixed 33% inflation plus encode and decode cost at both ends, on every chunk, in both directions. Opus over SRTP is compressed audio in a binary container — the comparison is not close, and on a metered mobile connection it is the difference a candidate notices.
Playback scheduling
Return audio is base64 PCM, so we decode each chunk into a Float32Array, build an AudioBuffer at the native output rate, and play it from a buffer source. Chunks must not overlap, so ours are strictly sequential: the next one starts when the current one fires onended.
source.onended = () => {
currentAudioSourceRef.current = null;
isPlayingAudioRef.current = false;
processAudioQueue(); // next chunk, immediately
};
source.start(0);
This is simple and it works, and it is also the weakest part of the design: every chunk boundary pays one event-loop hop. The better shape is to keep a running cursor and schedule each buffer at an absolute time on the audio clock, so the graph is always a few chunks ahead of the speaker. With WebRTC none of this exists as code, because the jitter buffer is doing it in C++.
One thing we did get right early: two AudioContext objects, one at the 16 kHz the model wants for input and one at the 24 kHz it emits. Matching both native rates deletes every resampling loop from the hot path. If you find yourself writing an interpolation loop per chunk in JavaScript, that is a sign your contexts are at the wrong rates.
The short version
WebRTC is not faster because UDP is faster. It is faster where it counts because the media stack it drags along — Opus, jitter buffer, loss concealment, congestion control — is exactly the code you would otherwise write badly in JavaScript, per chunk, on the main thread.
Audio-only negotiation, with the camera still on
Both rounds show the candidate their own video, and neither sends it anywhere. We request video and audio from getUserMedia, bind the stream to a local <video> element, and then add only the audio track to the peer connection.
const audioTrack = streamRef.current.getAudioTracks()[0];
audioSenderRef.current = pc.addTrack(audioTrack, streamRef.current);
// the video track is never added — it stays in the browser
This is worth doing deliberately. The camera feed drives the local presence signals and gives the candidate a mirror; uploading it would multiply bandwidth for zero conversational benefit, and it would turn a voice session into a video session in every sense that matters, including the privacy one. Audio-only negotiation is one line of restraint with a large payoff.
It also makes muting clean. Because we hold the RTCRtpSender, toggling the microphone is replaceTrack(null) rather than tearing down and renegotiating — and replaceTrack with a live track brings it back without a new offer. On the WebSocket side the equivalent is just not sending frames, which is easier still, but gives the far end no signal that you deliberately went quiet.
ICE, STUN and the TURN question
The part of WebRTC that surprises people who arrive from a WebSocket background is that connection establishment is a negotiation, not a handshake. We supply public STUN servers, the browser gathers candidates, and a candidate pair wins.
const pc = new RTCPeerConnection({
iceServers: [
{ urls: 'stun:stun.l.google.com:19302' },
{ urls: 'stun:stun1.l.google.com:19302' },
],
});
For a cloud voice endpoint, this is usually enough. You are not doing the hard NAT traversal case — two home routers trying to find each other — you are reaching a service with a public address, so a server-reflexive candidate does the job. That is why so many realtime voice integrations ship with two STUN URLs and no TURN server, and mostly get away with it.
Mostly. A network that blocks UDP outright cannot be fixed by any STUN server, because STUN only tells you your own mapped address. Corporate and campus networks do this, and it is exactly the population that shows up for interview practice from an office. The answer there is TURN over TCP or TLS on 443, which relays your media and therefore costs real bandwidth — so it is a deliberate, budgeted decision, not a default. What we do today is treat it as a connection failure and rebuild:
pc.oniceconnectionstatechange = () => {
if (pc.iceConnectionState === 'failed' ||
pc.iceConnectionState === 'disconnected') {
cleanupConnections();
setupWebRTCConnection();
}
};
A WebSocket has none of this. It is TCP on 443, indistinguishable from HTTPS to a firewall, and it connects essentially everywhere. On a hostile network, that is the WebSocket's single best argument — and it is a good one.
Interruption: who owns the audio already in flight
Barge-in is where the two designs stop resembling each other. An interviewer that keeps talking for two seconds after the candidate cuts in does not feel like a slow interviewer; it feels like a broken one.
Barge-in over WebRTC
We push turn detection to the server in the session configuration — semantic VAD, with interruption enabled — and then, mostly, do nothing. When the model decides the candidate has taken the floor, it stops sending audio. Because the only buffered audio lives in the browser's jitter buffer, measured in tens of milliseconds, the sound stops when the packets stop. The Realtime API guide covers the configuration; the point is that our client has no flush logic because there is nothing on our side to flush.
Barge-in over a WebSocket
Here we own the buffer, so we own the problem. The server tells us it detected an interruption, and that message is an instruction to throw away work we have already received and paid for:
if (serverContent.interrupted) {
audioQueueRef.current = []; // drop everything pending
if (currentAudioSourceRef.current) {
currentAudioSourceRef.current.stop(); // and cut the chunk mid-word
}
isPlayingAudioRef.current = false;
}
Miss either half and the bug is immediately audible. Clear the queue but forget to stop the playing source and the agent finishes its current chunk over the candidate. Stop the source but leave the queue and the next chunk starts the moment onended fires, which is worse — the agent appears to ignore the interruption entirely. This is the clearest illustration of the whole trade-off: the WebSocket design gave us full control over playback, and full control means the interruption semantics are now our bug surface. We wrote about the operational side of long-lived sockets in scaling real-time apps with WebSockets, and this is the application-level counterpart.
Reconnecting without spawning a second interview
Both transports drop. The failure modes are not symmetrical.
ICE failure is a state transition, so the WebRTC path watches the state machine and rebuilds the peer connection, as above. The subtlety is that a rebuild needs a valid credential, and ephemeral secrets expire — so an authentication error is not something to retry in a loop, it is a signal to send the candidate back for a fresh session.
A WebSocket close gives you a code, which is more information than it first appears. We only reconnect on an unexpected close, back off exponentially, and cap the attempts:
const shouldReconnect = event.code !== 1000 &&
reconnectAttempts < maxReconnects &&
!event.reason?.includes('expired');
if (shouldReconnect) {
const backoffDelay = 2000 * Math.pow(2, reconnectAttempts - 1);
reconnectTimeoutRef.current =
setTimeout(() => setupGeminiSession(true), backoffDelay);
}
Three rules are doing the work. Code 1000 is a clean close that we asked for, so reconnecting would fight our own teardown. An expired token will keep being expired, so retrying is guaranteed waste. And the attempt cap turns an unrecoverable network into one honest message instead of an infinite loop that quietly bills tokens. We also guard with a connecting flag and await the old socket's close before opening a new one — without that, a flaky connection produces two live sessions racing to answer the same candidate.
When each transport actually wins
- WebRTC wins when a human microphone is the source, when your users are on real-world networks you do not control, and when you would rather inherit a media stack than maintain one.
- WebSocket wins when the provider only offers a socket, when you need unencoded PCM for your own processing, when the audio originates on a server you control, or when UDP is blocked and you need something that looks like HTTPS.
- Both lose to a bad turn-detection configuration. No transport decision recovers a voice agent that interrupts the candidate mid-sentence or waits three seconds before answering.
Notice what is not on that list: raw milliseconds on a good connection. If you are choosing a transport to shave 50 ms off a pipeline whose model latency is measured in hundreds, you are optimising the wrong hop. Choose on tail behaviour instead — loss, interruption, reconnect — because that is where one transport keeps the conversation alive and the other quietly ruins it.
What we would pick starting over
WebRTC where the provider offers it, without hesitating. Not for latency on a good day, but because the interruption and loss behaviour we get for free is behaviour we demonstrably had to write, debug and re-debug on the socket path. Every bug in this article on the WebSocket side — the chunk that plays over an interruption, the double session on reconnect, the boundary gap between buffers — is a bug the media stack already solved.
And we would still keep one WebSocket round, because it is where the audio pipeline is visible. When you can see the 64 ms chunk, the base64 inflation and the playback queue in your own code, you understand what the other transport is doing on your behalf. If you want to hear both, the behavioural round and the coding round on Taqari are free — and if you are building something similar, our walkthrough of the OpenAI Realtime API over WebRTC and the notes on giving a browser voice agent context pick up where this one stops.
Frequently asked questions
Should a browser voice agent use WebRTC or WebSockets?
+
Use WebRTC when a person's microphone is the source. You inherit Opus encoding, a jitter buffer, packet-loss concealment and UDP transport. Reach for a WebSocket when the provider only offers one, when you need raw PCM, or when the audio originates on a server you control.
Is WebSocket audio actually slower than WebRTC?
+
Not on a clean network, where both are roughly one network hop. The gap opens under packet loss: TCP retransmits before releasing anything queued behind it, so one lost segment stalls the whole audio buffer, while WebRTC drops the frame and keeps talking.
Do you need a TURN server for a voice AI agent?
+
Often not. You are connecting to a cloud endpoint with a public address, so STUN is usually enough to get a candidate pair. You need TURN when a network blocks UDP outright — some corporate and campus networks do, and no STUN server can fix that.
Why capture audio at 16 kHz but play it back at 24 kHz?
+
Because those are the rates the model expects and produces. We run two AudioContexts at the native rates instead of one, which removes every resampling loop from the hot path. Resampling in JavaScript on each chunk is pure added latency and CPU.
How do you handle barge-in over a WebSocket?
+
You own it. When the server reports an interruption, empty your pending audio queue and stop the currently playing buffer source. Anything already decoded keeps playing otherwise, so the agent talks over the candidate for a second or two after they interrupt.
Does the transport choice matter next to model latency?
+
Less than people expect for the happy path, since model thinking time dominates. It matters enormously for the tail: degraded networks, interruptions and reconnects are where a transport either holds the conversation together or drops it.