Voice AI Platform Performance Analysis | SignalWire

Voice AI Platform Performance Analysis

Voice is fundamentally different from chat. In text interfaces, delays of several seconds are often tolerated. In live conversation, they are not.

1.24s

SW avg across 5 configs

38% faster than LiveKit tuned

1.09s

SW fastest config

1.75s

LiveKit tuned config

Key Findings

Faster in Every Config

SignalWire beats LiveKit and Vapi in every configuration tested. SignalWire Default with zero tuning (1.46s) still beats LiveKit's tuned config (1.75s) by 17%.

17-38% Faster (vs Tuned LiveKit)

Default config: 17% faster than LiveKit Tuned. SW Tuned: 37% faster. OSS models: 38% faster. Even our stock defaults beat their optimized configuration.

Orchestration Is the Variable

Same LLM, same TTS. The performance gap comes from the orchestration layer, not the models.

Fillers Cut Tool Call Silence 53%

Built-in speech fillers play during tool execution so the caller hears a response (1.40s) instead of silence (3.01s). Neither LiveKit nor Vapi offer equivalent filler support.

Vapi Dashboard vs Waveform

Vapi's dashboard displayed ~840ms. Stereo waveform analysis of actual calls measured 1.85s average. The dashboard may measure differently than what the caller experiences.

Consistency Over Averages

SignalWire's tightest config has a 0.09s spread (1.05-1.14s). LiveKit Tuned spread is 0.38s (1.61-1.99s). LiveKit baseline spread is 0.89s (1.55-2.44s).

Two Architectures, One Pipeline

Application-Layer (LiveKit, Vapi)

A Python async event loop orchestrates the pipeline from outside the media path

Audio goes to an external STT service, text streams to LLM, LLM output streams to an external TTS service, audio routes back to the caller

STT and TTS are separate network services. LLM-to-TTS streaming overlaps, but STT-to-LLM and TTS-to-caller are sequential hops

Barge-in is detected via VAD in the agent process, not at the raw audio level inside the media layer

Endpointing, turn detection, and tool execution all run in the same async loop as the pipeline orchestration

Getting it to work takes a weekend. Getting it to feel like a conversation takes months.

Latency variance comes from coordinating multiple external services, each with its own queue and network path

AI Kernel (SignalWire)

A purpose-built engine orchestrates the pipeline from inside the media processing layer, not from an application-layer script

The AI kernel calls the same external STT, LLM, and TTS services, but from a position with direct access to the audio stream, timing, and call state

The result is not fewer external calls. It is lower variance and faster coordination, because the orchestration layer was built for real-time audio from the start

Provider differences (latency profiles, streaming behavior, error modes) are normalized by the kernel so they perform consistently regardless of which vendor you choose

Barge-in and endpointing are detected at the audio level inside the media engine, before any external service is involved

Platform Comparison

Metric SignalWire LiveKit (tuned) Vapi
Conversational Latency 1.24s avg (5 configs) 1.75s 1.85s
Tool Turn Latency 2.01s (OSS+EL config) 2.90s 2.82s
Latency Spread 0.09-0.46s across configs 0.38s 0.66s
Filler Support Built-in, per-language and per-function Not available Partial, not fully controllable
OSS Model Support Drop-in via AI kernel OpenAI-compatible endpoints, Ollama helper Not tested
Endpointing Control Millisecond-level end_of_speech_timeout min/max delay (0.2s/0.8s tuned) Smart endpointing plans

Scaling Risk by Platform

Platform Avg Latency / Spread Primary Failure Mode at Scale
SignalWire 1.46s (default) / 0.46s — Low risk Narrow variance, few compounding failure boundaries. Predictable orchestration under concurrent load.
LiveKit (tuned) 1.75s / 0.38s — Low-Medium risk Tuning reduced spread from 0.89s to 0.38s. External STT/TTS dependencies remain.
LiveKit (baseline) 1.87s / 0.89s — Medium risk External STT/TTS dependencies multiply failure surface. Wider variance compounds under concurrent load.
Vapi 1.85s / 0.66s — Medium-High risk 4.87s outlier during a function call. Async tool execution is the likely cause.
Twilio 2.42s / 1.21s — High risk No programmatic control. No ability to tune, optimize, or instrument at scale. 3 of 4 turns exceeded 2.3s.

Human Conversational Thresholds

<2s: Responsive - Slight mechanical feel, but conversation flows naturally.

2-3s: Noticeable delay - Callers begin losing confidence in the system.

3s: Frequent interruption - Users frequently interrupt or assume the system has failed.

4s: Conversation breakdown - Callers often hang up or request a human.

Operational analytics at scale

The call_timeline feed provides every event from every call in a flat, queryable format. Export to Snowflake, BigQuery, Redshift, or any warehouse for operational analytics at scale.

Sigmond Runs on SignalWire

A voice and video AI agent built on the SignalWire SDK, wired to a live knowledge base, able to see what you show him. Build the same thing on voice, phone, WhatsApp, or SIP, or an entire white-label platform with SignalWire as the network underneath.

Test it yourself.

Sign up, deploy the same appointment agent, and measure your own latency. The tool is open source. The platform is pay-as-you-go. $0.16/min, no minimum.