ovos-wake-word-bench-mlsw-negatives-de-DE
收藏资源简介:
该数据集是用于 OVOS 插件竞技场(OVOS Plugin Arena)的样本集清单(manifest),并非音频语料库。它包含一个基于种子可复现的剪辑标识符列表,这些标识符选自 Multilingual Spoken Words Corpus(MLCommons, CC-BY-4.0),并记录了种子和源行数。该清单的目的是确保所有针对这些负样本进行评分的唤醒词插件都在完全相同的剪辑上运行,从而使不同插件之间的误接受率具有可比性。音频文件并未在此重新分发,用户需自行从原始来源获取。该清单采用 Apache-2.0 许可证。
This dataset is a manifest for the OVOS Plugin Arena, not an audio corpus. It contains a reproducible list of clip identifiers based on a seed, selected from the Multilingual Spoken Words Corpus (MLCommons, CC-BY-4.0), with records of the seed and source line numbers. The purpose of this manifest is to ensure that all wake word plugins scoring on these negative samples run on exactly the same clips, making false acceptance rates comparable across different plugins. Audio files are not redistributed here; users must obtain them from the original source. This manifest is licensed under Apache-2.0.
数据集概述
数据集名称:ovos-wake-word-bench-mlsw-negatives-de-DE
许可协议:Apache-2.0
语言:德语(de)
核心用途
这是为 OVOS 插件竞技场(Plugin Arena)构建的样本集清单(manifest),并非音频语料库本身。其主要用途是为德语唤醒词插件的性能评估提供统一的负面样本集合,确保不同插件在完全相同的音频片段上进行测试,从而保证误唤醒率(false-accept rates)的可比性。
数据来源与构建方式
- 源数据集:Multilingual Spoken Words Corpus(MLCommons 发布,采用 CC-BY-4.0 许可)
- 构建逻辑:通过固定的随机种子,从源数据集中可复现地选取一组片段标识符(clip identifiers),并记录种子信息及源数据行数。任何唤醒词插件在对抗这些负面样本进行评测时,都会使用完全相同的音频片段。
分发说明
- 音频文件并不在此仓库中重新分发,原始音频仍归 MLCommons 所有,遵循其本身的许可与署名条款。
- 本清单以 Apache-2.0 许可发布,与竞技场其他基准存储库保持一致。
相关链接
- 评测结果页面:https://openvoiceos.github.io/ovos-plugin-arena/
- 清单格式规范(
docs/SPECIFICATION.md):详见 https://github.com/OpenVoiceOS/ovos-plugin-arena





