ovos-wake-word-bench-mlsw-negatives-en-GB
收藏资源简介:
该数据集是一个用于唤醒词插件基准测试的负样本清单,名为ovos-wake-word-bench-mlsw-negatives-en-GB。它属于OVOS Plugin Arena项目,旨在提供一组可复现的剪片标识符列表,这些标识符从Multilingual Spoken Words Corpus(MLCommons, CC-BY-4.0)中通过固定种子和源行数选取。清单本身不包含音频数据,音频需从原始语料库获取。所有参与评估的唤醒词插件都使用相同的负样本集进行评分,从而确保误报率(false-accept rates)在不同插件之间具有可比性。清单采用Apache-2.0许可证,格式遵循OVOS Plugin Arena的SPECIFICATION文档。
This dataset is a negative samples manifest for wake word plugin benchmarking, named ovos-wake-word-bench-mlsw-negatives-en-GB. It belongs to the OVOS Plugin Arena project, providing a reproducible list of clip identifiers selected from the Multilingual Spoken Words Corpus (MLCommons, CC-BY-4.0) with a fixed seed and source row count. The manifest does not contain audio data; audio must be obtained from the original corpus. All participating wake word plugins use the same negative sample set for scoring, ensuring comparability of false-accept rates across different plugins. The manifest is licensed under Apache-2.0 and follows the SPECIFICATION document of OVOS Plugin Arena.
数据集概述
本数据集是 ovos-wake-word-bench-mlsw-negatives-en-GB,它不是一个音频语料库,而是一个样本集清单(manifest),专为 OVOS Plugin Arena 评测平台设计。
核心用途
- 提供一组可复现、基于固定随机种子的音频片段标识符列表
- 确保所有唤醒词插件在完全相同的负面样本上接受评测
- 使不同插件之间的**误唤醒率(false-accept rates)**具有可比性
数据来源与许可
- 音频片段选自 Multilingual Spoken Words Corpus(MLCommons,CC-BY-4.0 许可)
- 本数据集不重新分发音频,音频仍归属 MLCommons 及其原始许可条款
- 本清单本身以 Apache-2.0 许可发布,与竞技场其他基准仓库保持一致
记录信息
- 记录了随机种子和源数据行数,确保可复现性
- 语言范围标记为英文(en),地理区域为英国(GB)
关联链接
- 评测结果页面:https://openvoiceos.github.io/ovos-plugin-arena/
- 清单格式规范见:https://github.com/OpenVoiceOS/ovos-plugin-arena 中的
docs/SPECIFICATION.md




