Quantum Automations Quantum Automations
Blog · Portfolio
← Back to Blog
Guide · Voice AI

WebRTC vs SIP for UK Voice Agents: Stack Comparison

Published October 2026
Topic Voice Agents · Transport Architecture
Reading time 10 min
For UK SME ops leads
On this page
  1. How WebRTC and SIP transport layers actually differ under a live voice agent workload
  2. Latency breakdown: where each protocol adds milliseconds on UK calls
  3. Cost per minute at scale: elastic SIP vs carrier trunking for 10k monthly outbound calls
  4. Failover in production: SIP trunk failover vs WebRTC ICE reconnect under load
  5. When WebRTC wins and when SIP trunking wins: a decision framework
  6. Stack configurations: Retell + Twilio, VAPI + Vonage, and custom Asterisk setups
  7. UK number management: CLI presentation, Ofcom registration, and porting per stack
  8. What changed in 2025–2026: SIP trunking pricing and WebRTC-native agent orchestrators
  9. Good / Bad / Ugly: three transport decisions and their production consequences
  10. FAQ

We built the same qualification agent twice for a Bristol logistics firm — once on Twilio Programmable Voice, once with a direct SIP trunk into a co-located media server. Same LLM, same STT, same TTS. The Twilio build averaged 820ms from speech-end to first audio byte. The SIP trunk build averaged 640ms. On a 30-second qualification call, that 180ms gap produced two barge-in events that wouldn't have happened — the agent spoke over the caller as they drew breath between sentences. Twilio's media gateway adds 40–80ms of round-trip latency on UK calls compared to a direct SIP trunk, and when you're targeting sub-700ms total response time, that margin matters.

This post works through why that gap exists, what it costs in money and reliability at scale, and how to pick the right transport for your deployment.

How WebRTC and SIP transport layers actually differ under a live voice agent workload

SIP (Session Initiation Protocol) is a signalling protocol. It sets up the call, negotiates codecs, and tears it down — but the audio travels separately over RTP (Real-time Transport Protocol). A SIP trunk from a UK carrier delivers RTP packets directly to your media server with no browser stack in the path.

WebRTC wraps RTP in DTLS-SRTP for encryption and adds ICE (Interactive Connectivity Establishment) for NAT traversal. It was designed for browser-to-browser or browser-to-server calls where you don't control both endpoints. Twilio's Programmable Voice product uses a WebRTC gateway that handles the PSTN side and delivers WebRTC audio to your application endpoint — useful if you need a quick deployment, but it adds hops.

Under a voice agent workload, the difference shows in three places:

Media path length. A SIP trunk from a UK carrier routes audio through one or two hops inside the UK. A Twilio call routes through Twilio's media gateway, typically in Dublin or Frankfurt for UK numbers, before arriving at your media server.

Codec handling. SIP trunks between UK carriers typically use G.711 alaw end-to-end without transcoding. Twilio's WebRTC stream uses Opus, requiring a codec conversion step when your STT endpoint expects PCM or G.711.

Connection setup overhead. A DTLS handshake on WebRTC connection establishment adds 20–50ms at call start. For an outbound dialler placing hundreds of calls in a session, that overhead is consistent.

Neither protocol is inherently better. WebRTC wins on deployment simplicity. SIP wins on latency and cost at volume.

Latency breakdown: where each protocol adds milliseconds on UK calls

Here's how latency stacks for a typical UK outbound agent call, measured from speech-end-detection to first audio byte returned:

Component Twilio WebRTC Direct SIP Trunk
Network (PSTN to media server) 60–90ms 20–40ms
Codec transcoding 10–20ms 0ms (G.711 passthrough)
STT — Deepgram Nova-2, UK region 180–220ms 180–220ms
LLM — GPT-4o, streamed first token 200–280ms 200–280ms
TTS — ElevenLabs streaming 120–160ms 120–160ms
Total p50 570–770ms 520–700ms
Total p95 900–1,100ms 720–900ms

The STT, LLM, and TTS components are identical — they don't change with transport choice. What the transport choice controls is the floor beneath them. A direct SIP trunk lowers your p50 floor by 40–80ms and reduces p95 jitter because you've removed one gateway from the audio path.

For agents targeting sub-700ms total response time — a common threshold before interruption rates climb noticeably — that floor matters. For agents handling payment reminders or appointment confirmations where a few hundred milliseconds of latency doesn't trigger barge-in, Twilio's WebRTC stack is operationally simpler and entirely adequate.

Cost per minute at scale: elastic SIP vs carrier trunking for 10k monthly outbound calls

Twilio UK outbound (to landlines and mobiles) runs approximately £0.0085 per minute as of Q3 2026 using their elastic SIP product. No committed spend, no minimums. At 10,000 minutes per month, that's £85.

A carrier SIP trunk from Vonage or Telnyx costs £0.003–0.005 per minute for UK termination, but comes with fixed monthly overheads: DID rental at £2–4 per number, SIP trunk registration fees of £30–80 per month, and potentially a minimum monthly commitment. At 10,000 minutes with four active numbers and a standard trunk fee, you're looking at £55–65 per month — a saving of £20–30.

That saving doesn't justify the operational overhead below roughly 6,000–7,000 minutes per month. Above 20,000 minutes per month, the saving runs to £200–400 monthly and the case is clear.

One cost factor that catches teams out: mobile termination in the UK is priced separately from landline. Calls to 07xx numbers cost £0.012–0.016 per minute on Twilio, and carrier trunks charge similarly elevated rates for mobile. If your outbound list is predominantly mobile numbers — common in SME sales and logistics — model both scenarios with your actual ratio before switching.

Failover in production: SIP trunk failover vs WebRTC ICE reconnect under load

SIP trunk failover operates at the carrier routing level. You configure a primary and backup SIP URI on your trunk. If the primary stops responding to REGISTER or OPTIONS pings, the carrier routes to the backup. A working Asterisk PJSIP configuration looks like this:

; pjsip.conf — dual-trunk failover for UK outbound
[vonage-primary]
type=trunk
host=sip.vonage.com
port=5060
transport=udp
qualify_frequency=10
qualify_timeout=3.0

[vonage-backup]
type=trunk
host=sip2.vonage.com
port=5060
transport=udp
qualify_frequency=10
qualify_timeout=3.0

; dialplan: try primary, fall back on network failure
exten => _0X.,1,Dial(PJSIP/${EXTEN}@vonage-primary,,r)
 same => n,GotoIf($["${DIALSTATUS}" = "CONGESTION"]?fallback)
 same => n,GotoIf($["${DIALSTATUS}" = "CHANUNAVAIL"]?fallback)
 same => n,Hangup()
 same => n(fallback),Dial(PJSIP/${EXTEN}@vonage-backup,,r)
 same => n,Hangup()

With qualify_frequency=10, Asterisk sends OPTIONS pings every 10 seconds. Failover detection happens within 10–15 seconds — a mid-call interruption, not a silent handover. If you need sub-second failover, you need both trunks active simultaneously with load balancing, which adds complexity.

WebRTC ICE reconnect is a different mechanism. ICE handles NAT traversal for peer-to-peer calls — it wasn't designed as a carrier-level failover tool. If Twilio's media gateway drops, the WebRTC session attempts an ICE restart taking 5–30 seconds; the call usually drops. Twilio carries a 99.95% uptime SLA and redundant regional gateways, so this is an edge case — but worth accounting for in your incident response plan.

For most UK SME deployments, Twilio's reliability is sufficient. SIP trunk failover gives more vendor independence, but only if you've tested it end-to-end.

When WebRTC wins and when SIP trunking wins: a decision framework

Use Twilio WebRTC (or Retell/VAPI with default Twilio carrier) when: - Monthly volume is under 8,000 minutes - You're building fast and need to ship inside a week - Target response time is 700ms or above - You lack ops resource to manage SIP trunk configuration and monitoring

Use direct SIP trunking when: - Monthly volume exceeds 10,000 minutes - You're targeting sub-700ms response time for high-cadence conversation - You need number portability from a legacy telephony system - You're co-locating your media server in the same UK data centre as your carrier's point of presence

The hybrid approach: Run Twilio for early-stage outbound testing and SIP trunk for high-volume inbound. Inbound calls on your own SIP trunk give you more control over queue behaviour and call routing without the latency pressure that applies to real-time outbound conversation.

Stack configurations: Retell + Twilio, VAPI + Vonage, and custom Asterisk setups

Retell + Twilio is the default path: Retell handles agent orchestration, Twilio handles telephony. Audio travels via WebRTC from Twilio to Retell's media bridge. Retell's infrastructure is primarily US-West, adding 80–120ms to UK STT latency compared to a UK-hosted media server. Retell's custom telephony documentation covers replacing Twilio with a carrier SIP trunk while keeping Retell's agent orchestration layer intact.

VAPI + Vonage performs better for UK deployments. Vonage has SIP points of presence in London, and VAPI accepts SIP trunks directly. Configure the trunk in your VAPI outbound call request:

{
  "phoneNumberId": "your-vapi-number-id",
  "customer": {
    "number": "+441234567890"
  },
  "transport": {
    "provider": "vonage",
    "sipUri": "sip:[email protected]",
    "codec": "PCMA",
    "dtmfMode": "rfc2833"
  }
}

VAPI accepts a transport.sipUri override per call, so you can route different campaigns through different trunks without changing the agent configuration.

Custom Asterisk + carrier trunk gives maximum control and the lowest per-minute cost, at the price of full operational ownership. We built this configuration for a client running 40,000 outbound calls per month — the savings justified two days of initial configuration and ongoing monitoring. See our voice AI and document analysis build to understand the kind of deployment complexity that comes with a fully custom stack.

For a comparison of the orchestration platforms themselves, read Twilio vs Retell vs VAPI: which platform for UK voice agents.

UK number management: CLI presentation, Ofcom registration, and porting per stack

Ofcom's CLI presentation guidance requires every outbound call to present a dialable UK number. For voice agents making outbound calls, your presented CLI must be registered to your business, must not be withheld, and must connect to a person or voicemail when called back — not a disconnect tone or foreign number.

This affects your stack in three concrete ways:

Number registration. Twilio UK numbers are registered under Twilio's Ofcom allocation — you don't hold the registration directly. SIP trunk numbers from Vonage or BT Wholesale are registered to your business or the carrier on your behalf. The accountability chain is clearer if Ofcom queries CLI misuse.

Number porting. Moving a UK number from Twilio to a carrier SIP trunk takes 5–10 working days. Time your cutover away from peak outbound periods — we've seen dropped call volume during port windows when both carriers briefly claimed the number simultaneously.

CLI verification. Twilio enforces CLI matching — you can only present a number you own. On a raw SIP trunk you could technically present any CLI, which is a criminal offence under the Communications Act 2003. Ensure your SIP trunk provider enforces originating number verification, or add explicit CLI matching in your Asterisk dialplan.

For full detail on porting and registration procedures, read SIP trunking for UK voice agents: setup and cost breakdown.

What changed in 2025–2026: SIP trunking pricing and WebRTC-native agent orchestrators

Two developments shifted the calculus in the last 12 months.

First, UK carrier SIP trunking prices fell. Telnyx dropped their UK termination rate to £0.003 per minute in Q1 2026 following expansion of their London PoP. Vonage matched in Q2. That moved the cost crossover point down from roughly 12,000 minutes per month to 6,000–7,000 minutes, meaning direct SIP trunking now makes financial sense for a broader range of UK SME outbound deployments.

Second, WebRTC-native voice agent frameworks matured significantly. LiveKit Agents — originally a video conferencing infrastructure project — released a production-ready voice agent SDK in 2025 with WebRTC media processing built in. It handles STT, LLM integration, and TTS in the same session, eliminating one network round-trip compared to the Twilio-to-separate-media-server pattern. Early benchmarks show p50 latency of 480–520ms for UK deployments with Deepgram and ElevenLabs. That's competitive with a direct SIP trunk and comes with significantly less operational overhead. A fair counterpoint: webrtcHacks argues that for high-density outbound operations — diallers placing tens of thousands of concurrent calls — WebRTC's ICE negotiation and codec overhead still make carrier SIP the more predictable choice, regardless of SDK improvements. Both are defensible positions.

Good / Bad / Ugly: three transport decisions and their production consequences

Good: VAPI + Vonage SIP trunk for a UK SaaS renewal campaign

A 12,000-call-per-month renewal outbound campaign, mostly to 01/02 UK landlines. We configured VAPI with a Vonage SIP trunk, co-located our Deepgram STT endpoint in the same London facility as the Vonage PoP, and achieved consistent 560ms p50 latency. Cost dropped from £102/month on Twilio to £68/month on Vonage including trunk and DID fees. No porting issues — the client's numbers were already registered on Vonage.

Bad: Switching to carrier SIP mid-campaign without testing mobile termination separately

A Bristol outreach campaign had 70% mobile numbers on the list. We tested the new SIP trunk on landlines only. Mobile termination on the carrier we chose — a tier-2 UK operator — had an 8% call failure rate to certain MVNOs due to a gap in their roaming table. Twilio handles this silently with its own network fallbacks. We spent three days debugging what looked like a dialler concurrency issue before identifying the carrier. The fix was routing mobile calls back through Twilio and keeping only landlines on the SIP trunk. The latency saving disappeared for 70% of calls.

Ugly: Asterisk dual-trunk failover with the default OPTIONS interval

A client wanted the lowest possible cost per minute, so we built a custom Asterisk setup with two carrier SIP trunks and left the SIP OPTIONS keepalive interval at Asterisk's 60-second default. During a carrier maintenance window, the primary trunk failed silently. Asterisk took 60 seconds to detect it, dropping 40 active calls before the dialplan switched to backup. Fix: set qualify_frequency=10 in pjsip.conf. Failover now happens in 12–15 seconds — still a perceptible interruption, but not a 60-second blackout. For guidance on resilient dialplan design, see call-flow design for voice agents.

For latency improvements at the STT and LLM layers, read prompt engineering for voice agents.

FAQ

Can I switch from Twilio WebRTC to SIP trunking mid-project without rebuilding the agent logic?

Yes, in most cases. The transport layer sits below your agent logic — your STT, LLM, TTS, and call-flow code are unaffected. What changes is how audio reaches your media server: replace Twilio Programmable Voice with a SIP trunk pointing at the same media server, update your origination and termination config, and re-test end-to-end latency. Retell and VAPI both support custom SIP trunks as an alternative to their default carriers, so the agent orchestration code stays intact. Allow 5–10 working days if you need to port any UK numbers during the switch.

Which transport protocol does Retell.ai use by default, and can I swap it out?

Retell.ai defaults to Twilio Programmable Voice for telephony, which means your audio travels via WebRTC from Twilio's media gateway to Retell's media bridge. You can replace Twilio with a carrier SIP trunk by providing a SIP URI under Phone Numbers > Advanced in the Retell dashboard. Retell supports inbound SIP from carriers including Vonage and Telnyx. Outbound SIP trunking is configured via the phone_number resource in the Retell API, but requires a verified E.164 number on an approved carrier. The swap is worth doing above roughly 8,000–10,000 outbound minutes per month.

What Ofcom rules apply to CLI presentation on UK outbound AI calls?

Ofcom's CLI presentation requirements state that every outbound call must present a number the called party can use to call back or raise an objection. For AI voice agents, your presented CLI must be registered to your business, must not be withheld, and must connect to a live person or voicemail when dialled back — not a dead line or error tone. Presenting a non-geographic 08xx number as CLI on outbound AI calls is technically permitted but drives higher answer abandonment. A UK local number (01/02) performs better in answer-rate tests and is less likely to trigger carrier spam filters.

At what monthly call volume does carrier SIP trunking become cheaper than Twilio elastic SIP?

Twilio's UK outbound rate sits around £0.0085 per minute for standard Programmable Voice as of Q3 2026. Carrier SIP trunks from Vonage or Telnyx run £0.003–0.005 per minute at committed volume, plus fixed monthly overheads: DID rental at £2–4 per number and SIP trunk registration fees of £30–80 per month. Based on our modelling, a stack handling 10,000 outbound minutes per month breaks even at roughly 6,000–7,000 minutes. Above that volume the carrier trunk wins consistently. At 20,000 minutes per month the saving typically runs to £200–400 per month depending on your mobile versus landline termination mix.

Related Reading

Twilio vs Retell vs VAPI: Voice Agent Platform Comparison

An honest comparison of Twilio, Retell, and VAPI for voice agents: latency benchmarks, pricing, call-flow control, and a

SIP Trunking for UK Voice Agents: Setup and Cost

SIP trunking cuts per-minute costs 60% vs shared pools: UK carrier setup, number porting, and failover for production vo

Need a lower-latency voice agent stack for UK calls?

30-minute audit. We map your stack, your constraints, and where AI will pay back fastest.

Take the Quantum Leap →
© 2026 Quantum Automations Group Ltd
Home Blog Portfolio Privacy Terms Security