Voice Agent Simulation Bench (VAmoS) runs a credit card support agent through simulated phone calls and judges them on end-to-end task completion.
Leaderboard
| Voice agent | Task completion | Connect | Latency | Turns | Barge-in/call | Cost | STT | LLM | TTS |
|---|---|---|---|---|---|---|---|---|---|
| pipecat | 71.0% | 100.0% | 1.95s | 9 | 0.18 | ~$0.045 | Deepgram nova-3-general | gpt-4.1-mini | ElevenLabs |
| livekit | 70.3% | 100.0% | 2.37s | 7 | 0.07 | ~$0.048 | Deepgram nova-3-general | gpt-4.1-mini | ElevenLabs |
| grok voice | 69.3% | 99.0% | 1.93s | 7 | 0.05 | ~$0.129 | Grok Voice Think Fast 2.0 | ||
| vapi | 69.3% | 97.7% | 2.25s | 13 | 0.59 | ~$0.135 | Deepgram nova-3 | gpt-4.1-mini | ElevenLabs |
| elevenlabs | 67.3% | 99.7% | 1.19s | 7 | 0.27 | ~$0.114 | Scribe | gpt-4.1-mini | eleven_flash_v2 |
| gradium | 67.0% | 100.0% | 2.42s | 7 | 0.46 | ~$0.067 | Gradium STT | gpt-4.1-mini | Gradium TTS |
| cartesia | 67.0% | 100.0% | 2.23s | 7 | 0.02 | ~$0.146 | Ink | gpt-4.1-mini | Sonic |
| openai realtime | 67.0% | 100.0% | 1.53s | 7 | 0.15 | ~$0.074 | gpt-realtime-2 | ||
| hugging face | 65.3% | 100.0% | 5.39s | 5 | 2.86 | ~$0.155 | Whisper large-v3-turbo | DeepSeek V4 Flash | Kokoro-82M |
| deepgram | 64.0% | 99.7% | 2.06s | 7 | 0.66 | ~$0.112 | nova-3 | gpt-4.1-mini | aura-2-thalia-en |
| gemini 2.5 native | 64.0% | 99.3% | 15.95s | 7 | 0.01 | ~$0.023 | Gemini 2.5 Native Audio | ||
| gemini 3.1 live | 62.3% | 99.7% | 1.38s | 7 | 0.02 | ~$0.016 | Gemini 3.1 Flash Live | ||
| retell | 61.3% | 100.0% | 5.76s | 9 | 1.31 | ~$0.208 | Vendor ASR | gpt-4.1-mini | ElevenLabs |
| mistral small | 59.7% | 100.0% | 3.74s | 5 | 6.58 | ~$0.025 | Voxtral Realtime | mistral-small | Voxtral TTS |
| hermes | 58.7% | 100.0% | 10.43s | 5 | 1.32 | ~$0.069 | Deepgram nova-3 | hermes-4-70b | ElevenLabs |
| openai realtime mini | 51.3% | 96.3% | 1.79s | 9 | 0.57 | ~$0.032 | gpt-realtime-2.1-mini | ||
| nemotron* | 43.0% | 89.0% | 2.12s | 7 | 0.10 | n/a | Nemotron | Nemotron 3 Nano | Magpie |
Cost vs. completed tasks
Submit an implementation
This is a living benchmark. Every row is one implementation of Riley, the card-ops agent. Build your own on any framework, platform, or model, including ones already on the board. We’ll run it through the same 100 scenarios and publish the results. Feedback on the methodology is just as welcome.