A Soho restaurant group running 14 sites missed 31% of inbound calls during service hours. Front-of-house staff couldn't leave tables to answer, the phones rang out, and the bookings went to competitors. We deployed an inbound voice agent handling reservations, dietary queries, and cancellations across all 14 sites. Missed calls dropped to 4% in the first month at a cost equivalent to £1.80 per hour.
That £1.80 figure represents LLM inference, telephony, and TTS costs averaged across all 14 sites at peak volume — not a staff wage comparison. The marginal cost of answering one more call became negligible. A restaurant with a 70% fill rate and an average cover spend of £45 doesn't need many recovered bookings per week for the numbers to work.
Why hospitality inbound calls are worse than most sectors for AI handling
Restaurant calls are short, ambiguous, and arrive in conditions that break systems tuned on call-centre audio.
A caller saying "table for two Saturday" is missing a site, a time, a name, and a contact number. The agent has to extract all four without the conversation sounding like a form being read aloud. Compare that to dental or GP booking calls — a sector we've covered in our post on voice agents for dental booking and recall — where the call script is more predictable and the stakes for ambiguity are different.
The background noise problem is persistent. Calls during Friday and Saturday service windows come in with kitchen noise, ambient music at 75 dB, and callers standing on streets outside. We measured 18–22 dB of background noise on roughly 40% of calls during peak service at the Soho group. Most voice pipelines ship with default VAD thresholds tuned for quiet environments.
Dietary queries add a compliance layer absent from other booking categories. If a caller asks whether the lamb dish contains sesame, that is not a preference enquiry. It is a safety-critical question governed by the UK Food Information Regulations 2014. Getting that wrong isn't just a bad caller experience; it is a potential legal liability under food safety law.
Call-flow design for restaurant booking: the four call types that cover 94% of inbound volume
Before building anything, we ran call-type analysis on 3,400 inbound calls at the Soho group. Four types accounted for 94% of volume:
| Call type | Share of volume | Avg call length | Handling complexity |
|---|---|---|---|
| New reservation | 58% | 2 min 10 sec | Medium — slot lookup, party size, dietary note |
| Dietary / allergy query | 18% | 1 min 40 sec | High — allergen matrix, escalation logic |
| Cancellation or amendment | 12% | 1 min 55 sec | Medium — deposit rules, policy enforcement |
| General enquiry (parking, menu, hours) | 6% | 55 sec | Low — static knowledge base |
The remaining 6% were complaints, supplier calls, and delivery drivers — all routed to staff immediately. The agent handles the top four and escalates everything else. Don't build one agent that handles everything; build one that handles 94% correctly and design a clean escalation path for the rest. We documented the full call-flow node structure in our call flow design guide for voice agents.
Slot availability: connecting the voice agent to a live reservation system in real time
The most common failure mode in hospitality voice deployments isn't the language model — it's the reservation API integration. If the agent can't confirm slot availability during the call, the interaction fails. Callers won't accept "I'll check and someone will call you back"; they moved on from that pattern when online booking became standard.
The Soho group used ResDiary. Here's the reservation creation call the agent makes after extracting all required slots:
POST https://api.resdiary.com/api/v1/restaurants/{restaurantId}/reservations
Authorization: Bearer {session_token}
Content-Type: application/json
{
"partySize": 4,
"requestedDateTime": "2026-10-17T19:30:00",
"firstName": "Sarah",
"lastName": "Chen",
"contactPhone": "+447700900142",
"allergenNotes": "nut allergy — no nuts or nut derivatives",
"channelCode": "VOICE_AGENT"
}
The channelCode field is non-negotiable for attribution. You need to know how many covers came through the voice channel to evaluate whether the system is earning its keep. Without it, the covers are indistinguishable from web bookings.
API response latency from ResDiary averaged 320ms in our production logs. Combined with LLM inference (~400ms for GPT-4o mini at our call volume) and TTS synthesis (~180ms for ElevenLabs on cached phrases), total turn latency ran at roughly 900ms — just inside the threshold where callers don't read a pause as a technical fault. For a full breakdown of LLM choices for voice pipelines, see our LLM selection guide for voice agents.
Dietary and allergy queries: prompt design for safe, compliant responses
The allergen section of the system prompt is the most legally consequential part of the build. The agent needs to surface allergen information correctly and decline to venture into territory it isn't qualified to occupy.
Our approach limits the agent to three response types for allergen queries:
- "This dish contains [allergen]."
- "This dish may contain [allergen] due to preparation in a kitchen that also handles [allergen]."
- "I can't confirm allergen information for that dish — let me transfer you to our kitchen team who can give you a definitive answer."
The agent never says a dish is "safe" for someone with an allergy. Safety for a specific individual is a medical determination. The agent provides factual menu data only.
The system prompt includes a hard-coded escalation pattern: any mention of "anaphylaxis", "EpiPen", "severe reaction", or "anaphylactic" fires an immediate transfer to the duty manager. We added "anaphylactic" after noticing that callers used the adjective form in roughly 30% of relevant calls during testing.
The allergen matrix — a structured table mapping each dish to its declared and cross-contamination allergens — is injected into the system prompt as a JSON block at session initialisation. Kitchen staff update it via a simple CMS form; a nightly sync pushes the new version to the prompt template. The agent never queries an external allergen API mid-call. The data is in context at session start, which keeps latency predictable and eliminates a live dependency.
Cancellation and amendment handling: deposit rules and no-show policy enforcement
Cancellations are where deposit policy enforcement creates complexity. The agent needs to know whether the cancellation falls inside the refund window, what the refund amount is, and how to record the cancellation in the reservation system — and this varies by site and booking type.
Rather than encoding rules into the system prompt, we externalised deposit policy to a per-site config:
{
"site": "carnaby-street",
"deposit_policy": {
"standard": {
"refund_window_hours": 48,
"partial_refund_threshold_hours": 24,
"partial_refund_pct": 50
},
"events": {
"refund_window_hours": 168,
"partial_refund_threshold_hours": 72,
"partial_refund_pct": 0
}
}
}
The agent retrieves the applicable policy at call start, keyed on the DDI that was dialled. If a cancellation falls outside the refund window, the agent confirms the cancellation in the system but tells the caller the deposit is non-refundable per the booking terms, and flags the call for manager review if a dispute is likely.
No-show policy enforcement follows the same pattern. If a caller tries to cancel 30 minutes before a sitting to avoid a no-show fee, the agent still logs it as a customer cancellation — because it is. The policy decision about the fee belongs to the manager, not the agent.
Barge-in and interruption tuning for noisy restaurant environments
Barge-in sensitivity is a configuration problem disguised as a feature. In a quiet environment, you want the VAD threshold sensitive: callers interrupt for a reason and the agent should stop talking. In a loud restaurant environment, you want it less sensitive or background noise registers as barge-in and cuts off the agent mid-sentence.
We tuned the VAD threshold upward from the default — specifically to a dB threshold that filtered ambient noise under 60 dB. False barge-in triggers dropped from 34% of calls to 9% after the adjustment. The trade-off is that callers in particularly loud environments occasionally need to speak more deliberately; in our test corpus this caused problems in fewer than 2% of calls, which is an acceptable loss.
The full barge-in architecture — VAD sensitivity, suppression windows, and the interaction between ASR endpointing and LLM interruption handling — is covered in our guide to barge-in handling for voice agents.
Telephony setup: DDI routing, CID matching, and handoff to the duty manager
Each of the 14 sites got its own DDI. All 14 terminate on the same Twilio project, with routing logic that passes the dialled DDI as a metadata header to the voice agent. The agent uses that header to load the correct site configuration: menu knowledge, allergen matrix, deposit policy, and the site's duty manager DDI for transfers.
CID matching runs at call start. The Twilio webhook fires a lookup against the reservation database for the caller's number. If there's a confirmed booking in the next 48 hours, the agent opens with context: "Hi — I see you have a booking for four on Saturday. Are you calling about that reservation?" Roughly 22% of calls matched on CID, and those calls ran 35 seconds shorter on average than cold-start calls. At scale, that's a material reduction in telephony cost.
Transfer to the duty manager uses a warm SIP transfer. The agent delivers a spoken summary before dropping — caller name, party size, date, and any flagged dietary note — so the manager picks up with context rather than a cold "hello?". If the duty manager's DDI doesn't answer within 20 seconds, the call cascades to a site mobile and then deposits a voicemail with the transcript attached. See our voice agent transfer to human guide for the full transfer state machine.
Consult ICO guidance on automated processing before going live — particularly on whether booking confirmations constitute automated decisions under UK GDPR.
What changed in 2025–2026: real-time calendar APIs and LLM-native booking intent parsing
Two developments made hospitality voice agents substantially more practical over the past 18 months.
First, reservation platforms opened up their APIs. ResDiary's response times improved considerably; SevenRooms launched a public API tier in late 2025 that had previously required enterprise negotiation. OpenTable's Connect API now supports real-time availability lookups from third-party integrations — an integration path that was blocked or unreliable in 2023 and 2024.
Second, LLM-native intent parsing has replaced the rule-based slot-extraction pipelines that earlier builds depended on. GPT-4o and Claude 3.5 Sonnet handle multi-intent utterances — "I want to move my Saturday booking to Sunday and check whether you do a vegan option" — in a single inference pass. Earlier approaches required separate NER and intent classification steps that added latency and compounded failure rates, which is why single-site build times have dropped from 8–10 weeks to 3–4 weeks.
One counterpoint worth naming: research from Cornell's Centre for Hospitality Research has consistently found that a segment of hospitality guests prefers human contact at the booking stage, particularly for high-spend occasions or complex requirements. The voice agent doesn't replace the option of speaking to a person — it answers the calls that would otherwise go unanswered, while the transfer path is always available.
Good / Bad / Ugly: three hospitality deployments and what separated the good from the broken
Good — the Soho group (14 sites)
Missed calls: 31% to 4% in month one. The decisions that made it work: dedicated DDIs per site for clean routing and attribution, an allergen matrix maintained by kitchen staff rather than engineers, and a liberal escalation threshold. We set the agent to transfer freely on any ambiguous query rather than trying to resolve edge cases. Post-call surveys showed callers rated the experience as "normal" — the majority weren't certain whether they'd spoken to an automated system.
Bad — a single-site gastropub, Lancashire
The owner wanted to minimise staff involvement and set a very high intent-threshold before escalation. The agent tried to handle complex party-menu discussions and bespoke event enquiries that should have gone straight to the events coordinator. Complaint calls increased in the first two weeks. We rebuilt with a rule: anything involving more than 8 covers or a mention of "private dining" or "event" routes to a person immediately, regardless of intent confidence. Lesson: escalation thresholds need calibrating to the actual call mix for that site, not set once globally.
Ugly — a pub group with a legacy booking system
One site ran a booking system with no published API. We built an intermediate sync layer that polled the system's web interface and mirrored availability into a sidecar database the agent could query. It worked, but the refresh interval was 10 minutes, which meant the agent occasionally confirmed slots that had just been taken. Three double-bookings in the first month. We disclosed the limitation to the pub group and changed the agent's confirmation language to "provisional" with an SMS follow-up after the call. The lesson: if there's no API, the agent cannot give a real-time confirmation, and the call flow must be honest about that constraint rather than pretending otherwise.
For the document-analysis and knowledge-management side of voice agent builds, see our Voice AI and Document Analysis portfolio case study.