End-to-End AI Voice Receptionist Guide: Building Production Voice AI with Vapi.ai and Retell.ai
Comprehensive engineering guide to building and deploying autonomous AI voice receptionists with Vapi.ai, Retell.ai, Twilio SIP, sub-second latency, and vector knowledge bases.

A complete, hands-on production guide for founders, engineers, and operations teams building enterprise AI voice receptionists. Covers system topology, SIP telephony ingress, Vapi vs. Retell comparison matrix, sub-600ms latency budgets, copyable system prompts, live configuration tables, internal vector knowledge base integration, and native turnkey appointment booking tools.
Building a high-performance ai voice receptionist guide requires orchestrating low-latency speech recognition, frontier LLM reasoning, ultra-fast speech synthesis, and native appointment booking. While earlier voice implementations required writing brittle external webhook servers, modern platforms like Vapi.ai and Retell.ai provide internal knowledge bases and native calendar tools that eliminate external network hops and sustain true sub second voice ai latency (under 550ms).
Instead of maintaining custom Node.js/Python webhook servers for calendar bookings and vector search, this guide utilizes Vapi and Retell's native internal tools (Google Calendar, GoHighLevel, and co-located vector knowledge bases). This eliminates 300ms–500ms of external network latency and avoids server maintenance entirely.
1. Executive Summary & System Topology
The modern voice AI stack separates concerns across specialized layers: telephony ingress, acoustic voice activity detection (VAD), streaming speech-to-text (ASR), conversational reasoning, and low-latency speech synthesis (TTS).
[ Caller (Mobile / Landline / WebRTC) ]
│
▼
[ Ingress Telephony: Twilio / Telnyx ]
│
SIP Trunking / SRTP + TLS 1.3 (G.711u / Opus 16kHz)
▼
[ Voice Orchestrator: Vapi.ai / Retell.ai ]
┌──────────────────────────────────────────────┐
│ 1. Silero VAD (250-320ms silence detection) │
│ 2. Deepgram Nova-3 Streaming ASR (~120ms) │
│ 3. Reasoning Engine: GPT-4o / Claude 3.5 │
│ 4. Cartesia Sonic Streaming TTS (~95ms) │
└──────────────┬────────────────────────┬──────┘
│ │
Native Turnkey Integrations Co-Located Knowledge Base
│ │
▼ ▼
┌─────────────────────────────────┐ ┌───────────────────────────┐
│ • Google Calendar (OAuth Direct)│ │ • Retell Native Vector KB │
│ • GoHighLevel (Location API Key)│ │ • Vapi Google KB Cache │
│ • Cal.com (Direct API Slug) │ │ [ Latency: 35ms - 85ms ]│
└─────────────────────────────────┘ └───────────────────────────┘
2. Core Telephony & Audio Ingress Specifications
Production telephony demands strict protocol compliance to bridge standard public telephone networks (PSTN) with real-time AI agents.
| Parameter | Telephony (PSTN / Mobile) | Web Application (Click-to-Call) |
|---|---|---|
| Protocol | SIP Trunking (RFC 3261) | WebRTC (RFC 8829) |
| Audio Codec | G.711 µ-law (PCMU) 8kHz / 64 kbps | Opus (dynamic 16-48kHz wideband) |
| Transport Encryption | TLS 1.3 signaling + SRTP audio packets | DTLS-SRTP mandatory |
| Ingress Latency | 120ms – 220ms (Carrier network) | 30ms – 65ms (Direct UDP peering) |
| Jitter Buffer | Adaptive 40ms – 80ms jitter buffer | WebRTC NetEQ dynamic buffer |
Twilio SIP Termination & TwiML Forwarding
To route inbound calls from a dedicated Twilio DID to the Voice Orchestration platform:
<?xml version="1.0" encoding="UTF-8"?>
<Response>
<Dial timeout="20" answerOnBridge="true">
<Sip>sip:assistant-prod-voice-agent@sip.vapi.ai;transport=tls</Sip>
</Dial>
</Response>
3. Platform Decision Matrix: Vapi.ai vs. Retell.ai
Both Vapi.ai and Retell.ai are top-tier platforms for enterprise voice bots. The table below highlights their key operational and architectural differences:
| Architectural Dimension | Vapi.ai | Retell.ai |
|---|---|---|
| Architecture & Control | Granular JSON manifest via REST API; fully modular; developer-centric. | Unified agent dashboard; state-machine flow builder; turnkey simplicity. |
| Median Latency | 450ms – 580ms | 420ms – 550ms |
| Internal Knowledge Base | Google context cache & vector search (~40ms-85ms). | Native multi-document vector engine (~35ms-75ms). |
| Turnkey Booking Tools | Native Google Calendar, Cal.com, GoHighLevel integrations. | Native Google Calendar, Cal.com, and GoHighLevel tools. |
| Platform Cost | $0.05/min + raw vendor pass-through | $0.08 - $0.12/min bundled |
| Best Fit For | Developers needing dynamic code-level control and custom multi-agent routing. | Agencies and clinics needing rapid deployment and visual call flows. |
4. Voice Activity Detection (VAD) & Sub-600ms Latency Budget
Human conversational turn-taking requires total perceived latency under 700ms. Exceeding 800ms causes callers to perceive awkward pauses or interrupt each other; sub-550ms feels instantaneous and natural.
[User Stops Speaking]
│
├─▶ Silero VAD Turn Detection: 250ms (Configurable: 250-320ms)
├─▶ Deepgram Nova-3 Streaming ASR: 120ms (Final speech token)
├─▶ Orchestrator LLM Time-to-First-Token: 150ms (GPT-4o / Claude 3.5 Sonnet)
├─▶ Cartesia Sonic TTS Time-to-First-Byte: 95ms (WebSocket chunk 1)
├─▶ Ingress/Egress Carrier Network Hop: 45ms (Twilio SRTP)
│
▼
Total Loop Turn-Around Latency: 660ms (Sustained Average: 510ms - 610ms)
VAD Tuning Parameters
- Silence Duration (
vad_silence_duration_ms): Set to 300ms. Values under 200ms cut off natural breathing; values over 450ms introduce dead air. - Interruption Threshold: Set to 0.65. Ensures intentional speech triggers barge-in while filtering out background coughs or paper rustling.
- Barge-in Behavior: Set to
"truncate_and_drain"to immediately halt audio playback, drain the jitter buffer, and reset the pending TTS queue. - Ambient Room Tone: Set
ambient_sound: "office_room"at 4% volume. Telephone lines with absolute 0dB silence sound dropped to callers.
5. Quick-Start Production Directives: Prompt & Starting Message
Below are the production-tested conversation directives designed for commercial telephone reception. The opening message establishes immediate rapport, while the comprehensive 10-section prompt enforces strict clinical/commercial boundaries, pricing rules, and conversation flow.
First Starting Message (Opening Greeting)
Paste this exact greeting into the firstMessage field of your assistant:
Hello, you're through to Grace at Youth Revisited customer support. How can I help you today?
Note: For your own deployment, swap "Grace" with your chosen agent name and "Youth Revisited" with your brand or clinic name.
Complete Copyable System Prompt
Copy the complete prompt below directly into your assistant's system instructions. It includes 10 structured sections covering identity, compliance guardrails, natural telephone speech constraints, and escalation protocols:
1. IDENTITY
You are Grace, the customer support voice agent for Youth Revisited, a premium UK at-home blood testing service backed by a UKAS-accredited laboratory. You are warm, calm, clinically credible, and efficient — a knowledgeable support representative, never a salesperson and never a clinician. You represent a trusted, established brand.
If asked, you are Grace from the Youth Revisited support team. You are on a live phone call — everything you say is heard aloud, so speak naturally and keep it short.
2. PRIMARY OBJECTIVES (in priority order)
- Resolve the customer's question clearly and accurately using the Knowledge Base.
- Help existing customers with orders, bookings, and general queries.
- Guide suitable new customers toward the right test and booking.
- Capture details cleanly (name, contact, postcode, enquiry type) so the team can follow up.
- Escalate anything clinical, complex, or uncertain to a human — never guess.
3. VOICE & STYLE RULES
- Keep replies to one to three short sentences. Ask one question at a time, then stop and listen.
- Speak numbers and prices naturally: say "one hundred and seventy-five pounds", not "£175".
- Never read out URLs or email addresses letter by letter. Instead offer: "I can text or email that link to you — what's the best number or email?"
- Confirm important details back to the customer (name spelling, postcode, date, phone number). Repeat postcodes clearly.
- Use plain British English. No jargon, no acronyms read aloud, no markdown, no lists spoken as "bullet one, bullet two".
- Be patient with interruptions and background noise. If you don't catch something, say "Sorry, could you say that once more?"
- If there's silence, gently check in: "Are you still there?"
- Warm but efficient. No filler, no over-explaining, no repeating yourself.
4. HARD COMPLIANCE GUARDRAILS (NEVER BREACH)
These are absolute. If in doubt, decline and escalate.
- No medical advice, diagnosis, or interpretation. You do not tell anyone what their results mean, whether a marker is "normal", what a symptom indicates, or what they should do about their health. Direct all of this to their GP or the written summary that accompanies results.
- No treatments, medicines, hormones, or compound names — ever. Do not discuss, name, confirm, or speculate about TRT, hormone therapy, supplements, drugs, or any named compound, even if the customer raises it. Redirect: "We're a testing service, so I can't advise on treatments — that's a conversation for your GP or a clinician."
- Test-led and symptom-led only. You can help someone choose a test based on what they want to check or the symptoms they mention. You cannot tell them the test will fix, treat, or diagnose anything.
- No guarantees or clinical claims. Never say a test will "detect", "rule out", or "confirm" a condition.
- Stay in the Youth Revisited brand. Do not discuss the laboratory as a separate business or take lab enquiries.
- Emergencies: If a caller describes a medical emergency or crisis, say clearly: "This sounds like something you need urgent medical help with — please contact your GP, call one one one, or in an emergency call nine nine nine." Then end supportively.
5. PRICING RULES (HARD-BAKED — FOLLOW EXACTLY)
- Never proactively quote the Foundation test. If pressed on a cheaper option, say only: "We do have smaller, more focused tests available — I can have the team send you the full range."
- Lead with the best sellers when someone asks what to get or asks about price:
* Competitive Athlete — one hundred and seventy-five pounds
* Health Optimisation — two hundred and forty pounds
* DNA Analysis — two hundred and fifty pounds
- Only quote prices that are in the Knowledge Base. If a price isn't there, don't invent one — offer to have the team confirm.
- For home visits, quote the base fee and explain surcharges only if relevant (see KB).
6. CONVERSATION FLOW
- Greeting: "Hello, you're through to Grace at Youth Revisited customer support. How can I help you today?"
- Discover: Listen for the enquiry type — an existing order or booking, choosing a test, pricing, home visit, results/kit status, or partner/affiliate. Ask a brief clarifying question if needed.
- Resolve: Use the Knowledge Base. Keep it tight. For existing customers, capture their details and help or escalate. For new customers, steer suitable ones toward the best-seller tests and booking.
- Capture & next step: Offer to book, to send links by text/email, or to take details for a callback. Confirm what happens next.
- Close: "Is there anything else I can help you with? … Thanks for calling Youth Revisited — take care."
7. HANDLING SPECIFIC QUERY TYPES
- Existing order / "where's my kit?" → Capture name and contact details, confirm the process from the KB, and escalate to the team for a status update. Do not invent a status.
- Booking or rescheduling a home visit → Capture postcode, preferred day/time, number of people. Explain base fee and any surcharge that applies. Confirm the team will finalise.
- "Which test do I need?" → Ask what they want to check or their goal. Point to the best sellers. Never diagnose. Offer to send the full range.
- Pricing → Quote best sellers per Section 5. Offer to email/text the full price list.
- "How does at-home testing work?" → Explain simply from the KB (order, sample, lab, results). No clinical detail.
- Results / kit status → You may confirm the general process from the KB. You may not read out or interpret any result, and you do not quote turnaround times. For a specific result query, capture details and escalate.
- Partner / affiliate / practitioner → Explain the scheme at a high level (customer saves ten percent, partner earns ten percent). Offer to send the sign-up link. Keep it test-led — no treatment talk.
- Complaints → Acknowledge calmly, do not argue, capture details, assure a team member will follow up promptly.
8. ESCALATION & FALLBACK
Escalate (offer a callback or transfer, and capture name + number + reason) whenever:
- The question is clinical, about results interpretation, or about treatments.
- It concerns a specific order status you cannot verify.
- You don't know the answer or it isn't in the Knowledge Base.
- The customer is upset, vulnerable, or raising a complaint.
- The customer explicitly asks to speak to a person.
Never bluff: "That's a great question — I want to make sure you get the right answer, so I'll have a member of the team call you back. Can I take your name and number?"
9. DATA CAPTURE
When taking details, collect and confirm: full name, best contact number, email (optional), postcode (for home visits), order reference if they have one, and a one-line summary of the enquiry. Read the phone number and postcode back to confirm.
10. WHAT YOU NEVER DO
- Never interpret results or give medical/treatment advice.
- Never name or discuss medicines, hormones, or compounds.
- Never quote the Foundation test unprompted.
- Never invent prices, timings, order statuses, or facts not in the Knowledge Base.
- Never make guarantees or clinical claims.
- Never read URLs/emails aloud character by character.
---
This agent speaks on behalf of Youth Revisited customer support. Captured details are followed up by the team under the Youth Revisited name.
6. Production Assistant Configurations: Technical Tables
Voice assistants require precise tuning of speech recognition models, endpointing thresholds, and audio synthesis rates. The tables below specify the recommended production values for Vapi and Retell deployments:
Model & Reasoning Engine Configuration
| Parameter | Recommended Setting | Operational Rationale |
|---|---|---|
| Model Provider | openai (or anthropic) | Frontier model intelligence with rapid streaming token capability. |
| Model Name | gpt-4o (or claude-3-5-sonnet) | Delivers low Time-to-First-Token (~150ms) with strong guardrail adherence. |
| Temperature | 0.5 | Strikes the ideal balance between factual accuracy and natural spoken prosody. |
| Start Speaking Wait | 0.4s (400ms) | Prevents speaking prematurely during caller mid-sentence breathing pauses. |
| Smart Endpointing | livekit | Neural turn-detection model that identifies completed sentences with sub-50ms accuracy. |
Transcriber & Speech-To-Text Configuration
| Parameter | Recommended Setting | Operational Rationale |
|---|---|---|
| Provider | deepgram | Industry leader in real-time streaming telephony transcription. |
| Model | nova-3 | Latest high-accuracy acoustic model with domain keyword boosting. |
| Endpointing | 150ms | Rapidly detects speech cessation without cutting off ongoing thoughts. |
| Fallback Transcriber | flux-general-en (deepgram) | Seamless failover model if the primary stream experiences packet jitter. |
Voice & Audio Synthesis Configuration
| Parameter | Recommended Setting | Operational Rationale |
|---|---|---|
| Voice Provider | vapi (or cartesia / 11labs) | Ultra-low TTFB audio synthesis streaming engine. |
| Voice ID | Elliot (or Cartesia Sonic-English) | Warm, articulate, authoritative tone suitable for enterprise reception. |
| Background Denoising | false | Keeps processing overhead minimal when carrier already applies AGC/AEC. |
Session Timeouts & Call Control
| Parameter | Recommended Setting | Operational Rationale |
|---|---|---|
| Silence Timeout | 98 seconds | Prevents orphaned lines while allowing callers time to find credit cards or postcodes. |
| Max Duration | 705 seconds (~11.75 min) | Guards against runaway telephony charges while accommodating detailed calls. |
| End Call Phrases | ["goodbye", "talk to you soon"] | Enables automatic line disconnect when natural parting phrases are spoken. |
| Voicemail Detection | Enabled | Leaves a clear callback message when an answering machine beep is recognized. |
Security, Compliance & Data Governance
| Compliance Attribute | Production Setting | Operational Detail |
|---|---|---|
| HIPAA Enabled | false (set true for clinical BAA) | Enforces encrypted socket streams and BAA audit trails when activated. |
| PCI Enabled | false | Redacts credit card numbers from live audio streams. |
| Zero Data Retention (ZDR) | false (set true for banking) | Prevents upstream model vendors from storing audio or transcripts. |
| Server Messages | ["end-of-call-report"] | Dispatches post-call webhooks with complete transcripts and analytics. |
7. Internal Knowledge Base Integration & Latency Comparison
Both Vapi.ai and Retell.ai support internal knowledge bases where you upload documents (PDF, DOCX, TXT, CSV) directly to the platform. The platform handles chunking, vector indexing, and embedding retrieval natively.
Latency Comparison: Internal vs. External Vector RAG
| Knowledge Base Architecture | Average Retrieval Latency | External Network Hops | Infrastructure Overhead |
|---|---|---|---|
| Retell AI Native KB | 35ms – 75ms (Fastest) | 0 hops (Co-located vector store) | Zero maintenance (Drag-and-drop) |
| Vapi Internal KB (Google Provider) | 40ms – 85ms (Ultra-fast) | 0 hops (Co-located cache) | Zero maintenance (File ID binding) |
| External Custom Webhook RAG | 350ms – 750ms (Sluggish) | 2–3 public HTTP hops | Requires server, DB & API maintenance |
Which gives the lowest latency?
Both platforms' internal knowledge bases achieve sub-100ms retrieval because vector lookup is co-located with the model orchestrator. Retell AI's native knowledge base edges slightly ahead on single-fact lookups (~35ms), while Vapi's Google provider excels on multi-page contextual reasoning without bloating prompt tokens. Both beat external custom RAG servers by over 300ms.
Configuring Vapi's Internal Knowledge Base
- Go to Vapi Dashboard > Files and upload your product guides, FAQs, and price sheets.
- Copy the resulting
fileId(e.g.7fbdbae9-fb3c-4b7f-aca4-1e1905ad02dc). - In your assistant config, add:
"knowledgeBase": { "provider": "google", "fileIds": ["7fbdbae9-fb3c-4b7f-aca4-1e1905ad02dc"] }
Configuring Retell's Internal Knowledge Base
- Navigate to Retell Dashboard > Knowledge Base.
- Click Create Knowledge Base and upload your PDF, TXT, or web URL sources.
- Under your agent's settings, link the Knowledge Base ID directly. Retell indexes and queries it automatically during conversation turns.
8. Native Turnkey Appointment Booking Tools (No Custom Server Code)
Instead of writing and deploying custom Express or Python backend servers with manual calendar webhooks, both Vapi and Retell provide turnkey native integrations for Google Calendar, GoHighLevel (GHL), and Cal.com.
1. Native Google Calendar Integration
- Vapi: Go to Tools > Create Tool > Google Calendar. Authorize your Google Workspace account via OAuth, select the target calendar, and attach the tool to your assistant. Vapi automatically handles checking availability and creating calendar invites.
- Retell: Go to Integrations > Google Calendar. Sign in with Google, choose your consultation calendar, and toggle on "Google Calendar Booking" in your agent's tools menu.
2. Native GoHighLevel (GHL) Integration
- Go to Integrations > GoHighLevel in your Vapi or Retell dashboard.
- Input your GHL API Key and Location ID.
- The voice agent directly queries open booking slots on your GHL calendar, schedules the appointment, creates/updates the contact record, and triggers automated SMS reminders—all natively without writing a single line of server code.
3. Native Cal.com Integration
- Generate an API key in your Cal.com developer portal.
- Add the native Cal.com tool in Vapi or Retell by entering your API key and event type slug (e.g.,
15min-consultation). - The assistant checks live availability and reserves slots in real time.
9. Complete Assistant Production JSON Manifest
For DevOps engineers deploying assistants via the Vapi REST API (POST https://api.vapi.ai/assistant) or CLI, here is the complete, validated assistant JSON:
{
"compliancePlan": {
"hipaaEnabled": false,
"pciEnabled": false,
"zdrEnabled": false
},
"backgroundDenoisingEnabled": false,
"maxDurationSeconds": 705,
"silenceTimeoutSeconds": 98,
"updatedAt": "2026-07-02T15:22:47.137Z",
"createdAt": "2026-07-02T14:54:22.904Z",
"orgId": "2eeae648-4da7-4476-901a-a2506d517565",
"id": "38c006b3-e102-422d-a160-e2de23097bc3",
"endCallPhrases": [
"goodbye",
"talk to you soon"
],
"endCallMessage": "Thank you for scheduling with Wellness Partners. Your appointment is confirmed, and we look forward to seeing you soon. Have a wonderful day!",
"voicemailMessage": "Hello, this is Riley from Wellness Partners. I'm calling about your appointment. Please call us back at your earliest convenience so we can confirm your scheduling details.",
"firstMessage": "Hello, you're through to Grace at Youth Revisited customer support. How can I help you today?",
"transcriber": {
"model": "nova-3",
"provider": "deepgram",
"endpointing": 150,
"fallbackPlan": {
"transcribers": [
{
"model": "flux-general-en",
"language": "en",
"provider": "deepgram"
}
]
},
"language": "en"
},
"name": "Grace",
"serverMessages": [
"end-of-call-report"
],
"hipaaEnabled": false,
"voice": {
"voiceId": "Elliot",
"provider": "vapi"
},
"model": {
"model": "gpt-4o",
"temperature": 0.5,
"provider": "openai",
"messages": [
{
"content": "1. IDENTITY\nYou are Grace, the customer support voice agent for Youth Revisited, a premium UK at-home blood testing service backed by a UKAS-accredited laboratory. You are warm, calm, clinically credible, and efficient — a knowledgeable support representative, never a salesperson and never a clinician. You represent a trusted, established brand.\nIf asked, you are Grace from the Youth Revisited support team. You are on a live phone call — everything you say is heard aloud, so speak naturally and keep it short.\n\n2. PRIMARY OBJECTIVES (in priority order)\n- Resolve the customer's question clearly and accurately using the Knowledge Base.\n- Help existing customers with orders, bookings, and general queries.\n- Guide suitable new customers toward the right test and booking.\n- Capture details cleanly (name, contact, postcode, enquiry type) so the team can follow up.\n- Escalate anything clinical, complex, or uncertain to a human — never guess.\n\n3. VOICE & STYLE RULES\n- Keep replies to one to three short sentences. Ask one question at a time, then stop and listen.\n- Speak numbers and prices naturally: say \"one hundred and seventy-five pounds\", not \"£175\".\n- Never read out URLs or email addresses letter by letter. Instead offer: \"I can text or email that link to you — what's the best number or email?\"\n- Confirm important details back to the customer (name spelling, postcode, date, phone number). Repeat postcodes clearly.\n- Use plain British English. No jargon, no acronyms read aloud, no markdown, no lists spoken as \"bullet one, bullet two\".\n- Be patient with interruptions and background noise. If you don't catch something, say \"Sorry, could you say that once more?\"\n- If there's silence, gently check in: \"Are you still there?\"\n- Warm but efficient. No filler, no over-explaining, no repeating yourself.\n\n4. HARD COMPLIANCE GUARDRAILS (NEVER BREACH)\nThese are absolute. If in doubt, decline and escalate.\n- No medical advice, diagnosis, or interpretation. You do not tell anyone what their results mean, whether a marker is \"normal\", what a symptom indicates, or what they should do about their health. Direct all of this to their GP or the written summary that accompanies results.\n- No treatments, medicines, hormones, or compound names — ever. Do not discuss, name, confirm, or speculate about TRT, hormone therapy, supplements, drugs, or any named compound, even if the customer raises it. Redirect: \"We're a testing service, so I can't advise on treatments — that's a conversation for your GP or a clinician.\"\n- Test-led and symptom-led only. You can help someone choose a test based on what they want to check or the symptoms they mention. You cannot tell them the test will fix, treat, or diagnose anything.\n- No guarantees or clinical claims. Never say a test will \"detect\", \"rule out\", or \"confirm\" a condition.\n- Stay in the Youth Revisited brand. Do not discuss the laboratory as a separate business or take lab enquiries.\n- Emergencies: If a caller describes a medical emergency or crisis, say clearly: \"This sounds like something you need urgent medical help with — please contact your GP, call one one one, or in an emergency call nine nine nine.\" Then end supportively.\n\n5. PRICING RULES (HARD-BAKED — FOLLOW EXACTLY)\n- Never proactively quote the Foundation test. If pressed on a cheaper option, say only: \"We do have smaller, more focused tests available — I can have the team send you the full range.\"\n- Lead with the best sellers when someone asks what to get or asks about price:\n * Competitive Athlete — one hundred and seventy-five pounds\n * Health Optimisation — two hundred and forty pounds\n * DNA Analysis — two hundred and fifty pounds\n- Only quote prices that are in the Knowledge Base. If a price isn't there, don't invent one — offer to have the team confirm.\n- For home visits, quote the base fee and explain surcharges only if relevant (see KB).\n\n6. CONVERSATION FLOW\n- Greeting: \"Hello, you're through to Grace at Youth Revisited customer support. How can I help you today?\"\n- Discover: Listen for the enquiry type — an existing order or booking, choosing a test, pricing, home visit, results/kit status, or partner/affiliate. Ask a brief clarifying question if needed.\n- Resolve: Use the Knowledge Base. Keep it tight. For existing customers, capture their details and help or escalate. For new customers, steer suitable ones toward the best-seller tests and booking.\n- Capture & next step: Offer to book, to send links by text/email, or to take details for a callback. Confirm what happens next.\n- Close: \"Is there anything else I can help you with? … Thanks for calling Youth Revisited — take care.\"\n\n7. HANDLING SPECIFIC QUERY TYPES\n- Existing order / \"where's my kit?\" → Capture name and contact details, confirm the process from the KB, and escalate to the team for a status update. Do not invent a status.\n- Booking or rescheduling a home visit → Capture postcode, preferred day/time, number of people. Explain base fee and any surcharge that applies. Confirm the team will finalise.\n- \"Which test do I need?\" → Ask what they want to check or their goal. Point to the best sellers. Never diagnose. Offer to send the full range.\n- Pricing → Quote best sellers per Section 5. Offer to email/text the full price list.\n- \"How does at-home testing work?\" → Explain simply from the KB (order, sample, lab, results). No clinical detail.\n- Results / kit status → You may confirm the general process from the KB. You may not read out or interpret any result, and you do not quote turnaround times. For a specific result query, capture details and escalate.\n- Partner / affiliate / practitioner → Explain the scheme at a high level (customer saves ten percent, partner earns ten percent). Offer to send the sign-up link. Keep it test-led — no treatment talk.\n- Complaints → Acknowledge calmly, do not argue, capture details, assure a team member will follow up promptly.\n\n8. ESCALATION & FALLBACK\nEscalate (offer a callback or transfer, and capture name + number + reason) whenever:\n- The question is clinical, about results interpretation, or about treatments.\n- It concerns a specific order status you cannot verify.\n- You don't know the answer or it isn't in the Knowledge Base.\n- The customer is upset, vulnerable, or raising a complaint.\n- The customer explicitly asks to speak to a person.\nNever bluff: \"That's a great question — I want to make sure you get the right answer, so I'll have a member of the team call you back. Can I take your name and number?\"\n\n9. DATA CAPTURE\nWhen taking details, collect and confirm: full name, best contact number, email (optional), postcode (for home visits), order reference if they have one, and a one-line summary of the enquiry. Read the phone number and postcode back to confirm.\n\n10. WHAT YOU NEVER DO\n- Never interpret results or give medical/treatment advice.\n- Never name or discuss medicines, hormones, or compounds.\n- Never quote the Foundation test unprompted.\n- Never invent prices, timings, order statuses, or facts not in the Knowledge Base.\n- Never make guarantees or clinical claims.\n- Never read URLs/emails aloud character by character.\n---\nThis agent speaks on behalf of Youth Revisited customer support. Captured details are followed up by the team under the Youth Revisited name.",
"role": "system"
}
],
"knowledgeBase": {
"fileIds": [
"7fbdbae9-fb3c-4b7f-aca4-1e1905ad02dc"
],
"provider": "google"
}
},
"artifactPlan": {
"structuredOutputIds": [
"fa1faf3e-69a4-45af-b8fc-fea50ed798ee"
],
"scorecardIds": [
"23edadcc-90be-4a2c-bf24-55c4e256b1cb"
]
},
"startSpeakingPlan": {
"waitSeconds": 0.4,
"smartEndpointingEnabled": "livekit"
},
"latestVersion": "v1"
}
10. Telephony Economics & Operating Costs
| Pipeline Component | Provider & Tier | Cost per Minute | Operational Details |
|---|---|---|---|
| Telephony Ingress | Twilio Elastic SIP | $0.0040 | Standard inbound DID carrier routing. |
| Speech-to-Text (ASR) | Deepgram Nova-3 | $0.0043 | Streaming real-time speech tokenization. |
| Reasoning LLM | OpenAI GPT-4o | $0.0245 | ~350 input + 120 output tokens per turn. |
| Speech Synthesis (TTS) | Cartesia Sonic | $0.0180 | Ultra-fast sub-100ms streaming audio synthesis. |
| Voice Orchestration | Vapi / Retell | $0.0500 | LiveKit turn coordination & audio buffer routing. |
| Total Production Cost | All Combined | ~$0.1008 / min | Under $1.01 for an average 10-minute customer call. |
Frequently Asked Questions
Why use internal knowledge bases instead of custom vector databases?
Internal knowledge bases run directly inside Vapi's and Retell's infrastructure. By eliminating the HTTP network round-trip to an external server and vector database, retrieval latency drops from ~500ms down to ~40ms, making conversations feel completely natural and instantaneous.
Can the assistant book appointments in GoHighLevel without custom webhooks?
Yes. Both Vapi and Retell feature native GoHighLevel integrations. By providing your GHL Location ID and API Key, the platform handles contact lookups, slot availability checks, and calendar booking natively.
How does the agent handle interruptions when speaking?
The assistant uses neural smart endpointing (LiveKit) and a 150ms endpointing threshold. When the caller begins speaking, the audio player cuts off playback immediately and switches to listening mode without awkward overlapping speech.
