Ortem Technologies
    AI & Machine Learning

    Voice Agent Latency and Cost: Four Stacks Compared

    Praveen JhaAugust 29, 20268 min read
    Voice Agent Latency and Cost: Four Stacks Compared
    Quick Answer

    Across four production-shaped voice agent stacks on the same 1,000-minute support workload, the fastest reached 540 ms to first audio at p50 and 980 ms at p95 but cost the most at $0.19 per minute. The cheapest ran at $0.09 per minute with a weaker tail at 1,410 ms p95. The balanced enterprise stack (Deepgram, GPT, ElevenLabs, Twilio) hit 620 ms p50 and $0.14 per minute. Note that these are illustrative placeholder figures for editorial framing, not measured results.

    Commercial Expertise

    Need help with AI & Machine Learning?

    Ortem deploys dedicated AI & ML Engineering squads in 72 hours.

    Deploy Private AI

    Next Best Reads

    Continue your research on AI & Machine Learning

    These links are chosen to move readers from general education into service understanding, proof, and buying-context pages.

    Voice agent latency and cost comparison

    Voice agent comparisons usually quote vendor-published latency for a single component. That number does not survive contact with a real call, where speech-to-text, an LLM turn, tool calls and text-to-speech all sit in the path before the caller hears anything.

    This piece frames four full stacks measured end to end on the same workload.

    Read this first. Figures are illustrative placeholder data for editorial framing only. Replace with measured results from a controlled test run using the same network conditions, task script, and call volume.


    The workload

    A 1,000-minute mixed inbound and outbound support-call workload, run in the same network region, against the same evaluation script, using a production-shaped agent with speech-to-text, an LLM turn, tool calls and text-to-speech in the loop.

    The measurement is time to first audio — the delay between the caller finishing their turn and hearing the agent begin to respond. That is what a caller experiences as lag, and it is the number that determines whether a conversation feels natural.


    Results

    StackTime to first audio (p50)Time to first audio (p95)Cost per minute at volume
    A — Deepgram + GPT + ElevenLabs + Twilio620 ms1,120 ms$0.14/min
    B — Deepgram + Claude + ElevenLabs + Twilio710 ms1,280 ms$0.16/min
    C — Whisper RT + GPT Realtime + OpenAI TTS540 ms980 ms$0.19/min
    D — Open-source STT + local LLM + Cartesia + SIP780 ms1,410 ms$0.09/min

    Figure 1. In Ortem's illustrative voice-agent benchmark, the fastest stack reached 540 ms p50 to first audio, while the lowest-cost stack ran at $0.09 per minute at volume; the balanced enterprise stack delivered the best latency-cost tradeoff.

    Figures are illustrative placeholder data for editorial framing only. Replace with measured results from a controlled test run using the same network conditions, task script, and call volume.


    Where the stacks differ

    Stack C is fastest, and most expensive. Keeping the realtime path inside one vendor removes hand-off overhead, which shows in both p50 and the tail. Volume pricing is what you trade for it.

    Stack D is cheapest, with the weakest tail. At $0.09 per minute it is roughly half the cost of Stack C, but the 1,410 ms p95 means a meaningful share of turns feel slow. It also carries infrastructure management that the hosted options do not.

    Stack A is the balanced enterprise choice. 620 ms p50 at $0.14 per minute, with a tail under 1.2 seconds.

    Stack B trades latency for reasoning. 90 ms slower than Stack A at p50 for a different LLM in the loop — worth it when the workflow leans on tool use or multi-step reasoning, not worth it for scripted flows.


    Reading p95, not just p50

    The median is the number vendors quote and the tail is the number callers remember. Between Stack C and Stack D the p50 gap is 240 ms; the p95 gap is 430 ms. A stack that looks acceptable at the median can still produce a call where every fourth turn drags.

    If you are choosing on latency, set a p95 budget first and filter to the stacks that clear it, then compare cost among the survivors.


    Choosing on your own workload

    Cost per minute scales with volume, so the ranking above shifts at different call loads. Latency is dominated by whichever component is slowest in your specific region — a stack that wins in one region can lose in another.

    Both are reasons to measure your own path rather than adopt anyone's table, this one included.


    Building a voice agent and need the latency budget and stack choice worked out against your actual call patterns? Ortem Technologies' AI and ML solutions practice designs and ships production voice systems. See our guide to the best AI voice agents for the vendor landscape, or book a technical consultation →.

    About Ortem Technologies

    Ortem Technologies is a premier custom software, mobile app, and AI development company. We serve enterprise and startup clients across the USA, UK, Australia, Canada, and the Middle East. Our cross-industry expertise spans fintech, healthcare, and logistics, enabling us to deliver scalable, secure, and innovative digital solutions worldwide.

    📬

    Get the Ortem Tech Digest

    Monthly insights on AI, mobile, and software strategy - straight to your inbox. No spam, ever.

    voice agentsvoice AI latencyElevenLabsDeepgramCartesiaconversational AI costAI call center

    About the Author

    P
    Praveen Jha

    Director – AI Product Strategy, Development, Sales & Business Development, Ortem Technologies

    Praveen Jha is the Director of AI Product Strategy, Development, Sales & Business Development at Ortem Technologies. With deep expertise in technology consulting and enterprise sales, he helps businesses identify the right digital transformation strategies - from mobile and AI solutions to cloud-native platforms. He writes about technology adoption, business growth, and building software partnerships that deliver real ROI.

    Business DevelopmentTechnology ConsultingDigital Transformation
    LinkedIn

    Stay Ahead

    Get engineering insights in your inbox

    Practical guides on software development, AI, and cloud. No fluff — published when it's worth your time.

    Ready to Start Your Project?

    Let Ortem Technologies help you build innovative software solutions for your business.