Ortem Technologies
    AI & Machine Learning

    Voice Agent Latency and Cost: What Four AI Voice Stacks Delivered

    Ortem AI Research TeamAugust 26, 202611 min read
    Voice Agent Latency and Cost: What Four AI Voice Stacks Delivered
    Quick Answer

    In this illustrative test across four production-shaped voice stacks on the same 1,000-minute workload, the fastest reached 540 ms p50 time to first audio and 980 ms at p95 but cost the most at $0.19 per minute. The lowest-cost stack ran at $0.09 per minute with a weaker tail at 1,410 ms p95. The balanced enterprise stack, combining Deepgram, GPT, ElevenLabs and Twilio, delivered 620 ms p50 at $0.14 per minute. These are illustrative placeholder figures for editorial framing, not measured results.

    Commercial Expertise

    Need help with AI & Machine Learning?

    Ortem deploys dedicated AI & ML Engineering squads in 72 hours.

    Deploy Private AI

    Next Best Reads

    Continue your research on AI & Machine Learning

    These links are chosen to move readers from general education into service understanding, proof, and buying-context pages.

    For voice AI, latency is not a dashboard vanity metric. It determines whether a conversation feels natural or forces callers to wait, interrupt, repeat themselves, or request a human agent.

    The useful measure is not simply a vendor's claimed speech-to-text or text-to-speech speed. Teams need to measure the complete chain: telephony transport, speech recognition, turn detection, model inference, API or CRM lookup, speech synthesis, and audio delivery back to the caller.

    Read this first. Figures are illustrative placeholder data for editorial framing only. Replace with measured results from a controlled test run using the same network conditions, task script, and call volume.

    Short answer: In this illustrative test, the fastest stack reached 540 ms p50 time to first audio, while the lowest-cost stack operated at $0.09 per minute. The strongest enterprise balance came from a modular Deepgram, GPT, ElevenLabs, and Twilio architecture at $0.14 per minute.

    Ortem Technologies builds AI solutions, cloud systems, backend platforms, and customer-facing digital products. Our ClearVoice case study describes a financial-services voice AI agent using Twilio, Deepgram, ElevenLabs, Salesforce, and a core-banking API integration.


    Test conditions

    The benchmark used four representative production-style stacks against one fixed script: caller authentication, balance or policy lookup, FAQ response, appointment scheduling, and a handoff scenario.

    ControlTest condition
    Call volume1,000 combined inbound and outbound minutes
    NetworkSame telephony region and stable business-grade network path
    ScriptSame caller utterances, pauses, interruptions, API lookup sequence, and escalation rule
    MeasurementTime from caller end-of-turn to first audible agent response
    CostBlended telephony, STT, LLM, TTS, and infrastructure cost per call minute
    ExclusionsOne-time build, integration, and optimisation costs are reported separately

    Both p50 and p95 should be published. P50 shows the typical caller experience; p95 shows the slow tail that can damage perceived quality during peak load, retries, or slow external API calls.


    Voice agent latency and cost

    StackCore componentsTime to first audio p50Time to first audio p95Blended cost at volume
    Stack ADeepgram + GPT + ElevenLabs + Twilio620 ms1,120 ms$0.14/min
    Stack BDeepgram + Claude + ElevenLabs + Twilio710 ms1,280 ms$0.16/min
    Stack CWhisper RT + GPT Realtime + OpenAI TTS540 ms980 ms$0.19/min
    Stack DOpen-source STT + local LLM + Cartesia + SIP780 ms1,410 ms$0.09/min

    Figure 1. Illustrative controlled benchmark: Stack C produced the fastest median first-audio response, while Stack D produced the lowest per-minute operating cost. Stack A offered the most balanced latency-cost profile.

    Figures are illustrative placeholder data for editorial framing only. Replace with measured results from a controlled test run using the same network conditions, task script, and call volume.


    How to interpret the results

    Stack A: balanced enterprise deployment. At 620 ms p50 and $0.14 per minute, Stack A is appropriate for customer support, booking, account enquiries, and workflows requiring external data lookups. It avoids over-optimising for a headline latency number at the expense of reliability, observability, and integration flexibility.

    Stack B: higher reasoning allowance. This design incurred a slightly slower response profile because model and tool-routing decisions were more involved. It is more suitable where the agent must interpret policy language, apply complex decision rules, or make multi-step backend requests.

    Stack C: speed-first conversational experience. This stack gave the fastest p50 and p95 result. The trade-off is higher per-minute spend, making it better suited to premium voice experiences, short calls, high-value conversion flows, or applications where responsiveness directly affects completion.

    Stack D: cost-optimised architecture. Running more components under direct infrastructure control reduced call-minute cost. However, it produced the slowest p95 tail and carried a greater platform-management burden: GPU capacity, model updates, observability, failover, and incident response.


    The cost formula to publish

    Use a transparent blended calculation: cost per minute equals telephony plus speech-to-text plus LLM plus text-to-speech plus infrastructure.

    For a 10,000-minute monthly programme, the illustrative operating cost would range from $900 per month for Stack D to $1,900 per month for Stack C. This comparison excludes implementation, integration, security, monitoring, and human-supervision costs.


    How to improve voice-agent latency

    • Stream audio input and speech output rather than waiting for a full response.
    • Place voice, model, and application infrastructure in compatible regions.
    • Keep system prompts, retrieval context, and tool descriptions concise.
    • Return a short acknowledgement while a longer backend request completes.
    • Cache predictable answers and session information where appropriate.
    • Instrument each stage separately: STT, LLM first token, tool calls, TTS first audio, and telephony transport.
    • Design a clean human handoff for complex or low-confidence scenarios.

    Building a voice agent and need the latency budget and stack choice worked out against your actual call patterns? Ortem Technologies' AI and ML solutions practice designs and ships production voice systems. See our guide to the best AI voice agents, or book a technical consultation →.

    About Ortem Technologies

    Ortem Technologies is a premier custom software, mobile app, and AI development company. We serve enterprise and startup clients across the USA, UK, Australia, Canada, and the Middle East. Our cross-industry expertise spans fintech, healthcare, and logistics, enabling us to deliver scalable, secure, and innovative digital solutions worldwide.

    📬

    Get the Ortem Tech Digest

    Monthly insights on AI, mobile, and software strategy - straight to your inbox. No spam, ever.

    voice agentsvoice AI latencyElevenLabsDeepgramCartesiaconversational AI costAI call center

    About the Author

    O
    Ortem AI Research Team

    Technology Division, Ortem Technologies

    The Ortem AI Research Team is a cross-functional group of ML engineers, data scientists, and software architects embedded across our product, platform, and client delivery divisions. The team researches and evaluates emerging technologies — including large language models, agentic AI systems, computer vision, and MLOps infrastructure — translating complex concepts into actionable guidance for engineering leaders and enterprise decision-makers. Each article published under this byline is the result of collaborative investigation: real-world experimentation, architecture reviews, and performance benchmarking drawn from live client projects and internal R&D initiatives. The team is committed to publishing technically rigorous, vendor-neutral content that helps organisations cut through AI hype and make confident, ROI-driven technology decisions.

    Artificial IntelligenceMachine LearningData Engineering

    Frequently Asked Questions

    Stay Ahead

    Get engineering insights in your inbox

    Practical guides on software development, AI, and cloud. No fluff — published when it's worth your time.

    Ready to Start Your Project?

    Let Ortem Technologies help you build innovative software solutions for your business.