bosonai/IHBench
收藏资源简介:
IHBench是一个用于评估语音代理在结构化多步骤工作流中中断后恢复能力的基准测试。与衡量中断时机(如打断检测、端点检测、轮转)的基准不同,IHBench专注于评估中断后代理的响应:是否在正确步骤恢复工作流、处理用户插话,并避免重复用户已听过的内容。该基准包含45个合成生成且经过验证的对话,覆盖10个企业领域,具有428个中断点,涵盖六种中断类型(正常、不耐烦、纠正、话题切换、填充语、反驳)。每个中断都带有单独的中断评估标准,并在任务完成度和恢复质量两个轴上进行评分。数据集中每个用户回合的音频都直接嵌入。
IHBench evaluates post-interruption recovery in voice agents executing structured, multi-step workflows. Unlike benchmarks that measure the timing of interruptions (barge-in detection, endpointing, turn-taking), IHBench measures what the agent says after an interruption: does it resume the workflow at the correct step, address the users interjection, and avoid re-delivering content the user already heard? The benchmark contains 45 synthetically generated, verified conversations across 10 enterprise domains, with 428 interruption points spanning six interruption types (normal, impatient, correction, topic switch, filler, pushback). Each interruption carries a per-interruption evaluation rubric and is scored on two axes: task fulfillment and recovery quality. Audio for every user turn is embedded directly in the dataset.




