Kalyvox Voice Benchmark 2026
收藏资源简介:
Benchmark dataset measuring AI voice agent performance across 240 controlled test calls covering 12 standardized scenario families in French and English. The benchmark includes 120 French and 120 English calls conducted between August 22 and September 22, 2026. The dataset contains one row per controlled test call, including scenario information, expected and detected intent, response latency, intent correctness, task completion, transfer results, appointment booking results, fallback handling and call duration. Key benchmark results:- Median response latency: 735 ms- P95 response latency: 1,079 ms- Intent accuracy: 94.6%- Task completion: 91.2%- Successful call transfers: 96.7% (58/60)- Successful appointment bookings: 87.5% (35/40)- Correct fallback handling: 91.7% (55/60) Tests were performed using standardized scenarios designed to evaluate common AI receptionist tasks and failure cases. Full methodology, metric definitions and analysis:https://kalyvox.ai/en/ai-voice-agent-benchmark



