遇见数据集

Kalyvox Voice Benchmark 2026

收藏
Zenodo2026-09-27 更新2026-10-01 收录
官方服务:

资源简介:

Benchmark dataset measuring AI voice agent performance across 240 controlled test calls covering 12 standardized scenario families in French and English. The benchmark includes 120 French and 120 English calls conducted between August 22 and September 22, 2026. The dataset contains one row per controlled test call, including scenario information, expected and detected intent, response latency, intent correctness, task completion, transfer results, appointment booking results, fallback handling and call duration. Key benchmark results:- Median response latency: 735 ms- P95 response latency: 1,079 ms- Intent accuracy: 94.6%- Task completion: 91.2%- Successful call transfers: 96.7% (58/60)- Successful appointment bookings: 87.5% (35/40)- Correct fallback handling: 91.7% (55/60) Tests were performed using standardized scenarios designed to evaluate common AI receptionist tasks and failure cases. Full methodology, metric definitions and analysis:https://kalyvox.ai/en/ai-voice-agent-benchmark

提供机构:
Zenodo
创建时间:
2026-09-27
二维码
社区交流群
二维码
科研交流群
商业服务