Voice Agent Latency and Cost: Four Stacks Compared
Across four production-shaped voice agent stacks on the same 1,000-minute support workload, the fastest reached 540 ms to first audio at p50 and 980 ms at p95 but cost the most at $0.19 per minute. The cheapest ran at $0.09 per minute with a weaker tail at 1,410 ms p95. The balanced enterprise stack (Deepgram, GPT, ElevenLabs, Twilio) hit 620 ms p50 and $0.14 per minute. Note that these are illustrative placeholder figures for editorial framing, not measured results.
Commercial Expertise
Need help with AI & Machine Learning?
Ortem deploys dedicated AI & ML Engineering squads in 72 hours.
Next Best Reads
Continue your research on AI & Machine Learning
These links are chosen to move readers from general education into service understanding, proof, and buying-context pages.
AI & ML Solutions
Move from concept articles to real implementation planning for copilots, RAG, automation, and analytics.
Explore AI servicesAI Agent Development
See how Ortem builds autonomous workflows, tool-using agents, and human-in-the-loop systems.
View agent serviceAI Product Case Study
Study a production AI platform with architecture, launch scope, and operating model context.
Read case studyVoice agent comparisons usually quote vendor-published latency for a single component. That number does not survive contact with a real call, where speech-to-text, an LLM turn, tool calls and text-to-speech all sit in the path before the caller hears anything.
This piece frames four full stacks measured end to end on the same workload.
Read this first. Figures are illustrative placeholder data for editorial framing only. Replace with measured results from a controlled test run using the same network conditions, task script, and call volume.
The workload
A 1,000-minute mixed inbound and outbound support-call workload, run in the same network region, against the same evaluation script, using a production-shaped agent with speech-to-text, an LLM turn, tool calls and text-to-speech in the loop.
The measurement is time to first audio — the delay between the caller finishing their turn and hearing the agent begin to respond. That is what a caller experiences as lag, and it is the number that determines whether a conversation feels natural.
Results
| Stack | Time to first audio (p50) | Time to first audio (p95) | Cost per minute at volume |
|---|---|---|---|
| A — Deepgram + GPT + ElevenLabs + Twilio | 620 ms | 1,120 ms | $0.14/min |
| B — Deepgram + Claude + ElevenLabs + Twilio | 710 ms | 1,280 ms | $0.16/min |
| C — Whisper RT + GPT Realtime + OpenAI TTS | 540 ms | 980 ms | $0.19/min |
| D — Open-source STT + local LLM + Cartesia + SIP | 780 ms | 1,410 ms | $0.09/min |
Figure 1. In Ortem's illustrative voice-agent benchmark, the fastest stack reached 540 ms p50 to first audio, while the lowest-cost stack ran at $0.09 per minute at volume; the balanced enterprise stack delivered the best latency-cost tradeoff.
Figures are illustrative placeholder data for editorial framing only. Replace with measured results from a controlled test run using the same network conditions, task script, and call volume.
Where the stacks differ
Stack C is fastest, and most expensive. Keeping the realtime path inside one vendor removes hand-off overhead, which shows in both p50 and the tail. Volume pricing is what you trade for it.
Stack D is cheapest, with the weakest tail. At $0.09 per minute it is roughly half the cost of Stack C, but the 1,410 ms p95 means a meaningful share of turns feel slow. It also carries infrastructure management that the hosted options do not.
Stack A is the balanced enterprise choice. 620 ms p50 at $0.14 per minute, with a tail under 1.2 seconds.
Stack B trades latency for reasoning. 90 ms slower than Stack A at p50 for a different LLM in the loop — worth it when the workflow leans on tool use or multi-step reasoning, not worth it for scripted flows.
Reading p95, not just p50
The median is the number vendors quote and the tail is the number callers remember. Between Stack C and Stack D the p50 gap is 240 ms; the p95 gap is 430 ms. A stack that looks acceptable at the median can still produce a call where every fourth turn drags.
If you are choosing on latency, set a p95 budget first and filter to the stacks that clear it, then compare cost among the survivors.
Choosing on your own workload
Cost per minute scales with volume, so the ranking above shifts at different call loads. Latency is dominated by whichever component is slowest in your specific region — a stack that wins in one region can lose in another.
Both are reasons to measure your own path rather than adopt anyone's table, this one included.
Building a voice agent and need the latency budget and stack choice worked out against your actual call patterns? Ortem Technologies' AI and ML solutions practice designs and ships production voice systems. See our guide to the best AI voice agents for the vendor landscape, or book a technical consultation →.
About Ortem Technologies
Ortem Technologies is a premier custom software, mobile app, and AI development company. We serve enterprise and startup clients across the USA, UK, Australia, Canada, and the Middle East. Our cross-industry expertise spans fintech, healthcare, and logistics, enabling us to deliver scalable, secure, and innovative digital solutions worldwide.
Get the Ortem Tech Digest
Monthly insights on AI, mobile, and software strategy - straight to your inbox. No spam, ever.
About the Author
Director – AI Product Strategy, Development, Sales & Business Development, Ortem Technologies
Praveen Jha is the Director of AI Product Strategy, Development, Sales & Business Development at Ortem Technologies. With deep expertise in technology consulting and enterprise sales, he helps businesses identify the right digital transformation strategies - from mobile and AI solutions to cloud-native platforms. He writes about technology adoption, business growth, and building software partnerships that deliver real ROI.
Stay Ahead
Get engineering insights in your inbox
Practical guides on software development, AI, and cloud. No fluff — published when it's worth your time.
Ready to Start Your Project?
Let Ortem Technologies help you build innovative software solutions for your business.
You Might Also Like
What an Engineering Org Actually Spends on AI Tooling

AI App Development Cost in 2026: Real Numbers from Shipped Projects

