livekit/eot-bench-data
收藏资源简介:
该数据集是LiveKit端到端转捩检测基准测试的一个最小公共子集,从完整发布就绪数据集中采样而来。采样是确定性的:使用种子20260603对每种语言最多采样400个turn(无放回)。如果某种语言的可用行数少于请求的样本大小,则包含所有可用行(例如:阿拉伯语381行、印尼语396行、日语356行、韩语377行、中文344行)。数据集包含多种语言配置(如英语、德语、日语等)和一个all配置(所有可用采样语言)。所有配置使用相同的数据模式:包括全局唯一的turn ID(id)、16 kHz单声道WAV音频(audio)、标准化语言代码(language)、音频持续时间(duration)、候选沉默时间跨度(silence_spans)、词级转录时间戳(words)和先前的对话消息(messages)。数据集中排除了最终沉默间隔短于0.2秒的turn(以便在0.2秒评分点进行评估),也排除了任何沉默间隔长于5.0秒的turn(以使基准测试专注于普通的端点检测停顿)。数据集支持多种语言,包括英语、阿拉伯语、德语、西班牙语、法语、印地语、印尼语、意大利语、日语、韩语、荷兰语、葡萄牙语、土耳其语和中文,总行数为5454。
This dataset contains a minimal public subset of the LiveKit end-of-turn benchmark. It is sampled from the full release-ready dataset preserved on the full revision. The sample is deterministic: up to 400 turns per language were sampled without replacement using seed 20260603. Languages with fewer than the requested sample size include all available rows: ar (381 available), id (396 available), ja (356 available), ko (377 available), zh (344 available). Configs include all (all available sampled languages) and language-specific configs (e.g., en, de, ja). All configs use the same schema: id (globally unique turn id), audio (16 kHz mono WAV audio), language (normalized language code), duration (audio duration in seconds), silence_spans (candidate silence spans with start and end), words (word-level transcript timings), messages (prior conversation messages). Rows whose final end-of-turn silence span is shorter than 0.2s are excluded so every turn can be evaluated at a 0.2s score point, and rows with any silence span longer than 5.0s are excluded to focus on ordinary endpointing pauses. Languages supported: English, Arabic, German, Spanish, French, Hindi, Indonesian, Italian, Japanese, Korean, Dutch, Portuguese, Turkish, and Chinese, with a total of 5454 rows.




