ovos-wake-word-bench-mlsw-negatives-nl-NL
收藏资源简介:
该数据集是用于OVOS插件竞技场唤醒词基准测试的样本清单,并非音频语料库。它包含从Multilingual Spoken Words Corpus(MLCommons, CC-BY-4.0)中选取的荷兰语(nl-NL)剪辑标识符列表,并附有种子和源行计数,以确保可重现性。音频文件未在此分发,需从MLCommons获取。该清单用于评估唤醒词插件,使不同插件之间的误接受率具有可比性。数据集以Apache-2.0许可证发布。
This dataset is a sample list for the OVOS Plugins Arena wake word benchmark, not an audio corpus. It contains a list of Dutch (nl-NL) clip identifiers selected from the Multilingual Spoken Words Corpus (MLCommons, CC-BY-4.0), along with seed and source line counts to ensure reproducibility. The audio files are not distributed here and must be obtained from MLCommons. This list is used to evaluate wake word plugins, making false acceptance rates comparable across different plugins. The dataset is released under the Apache-2.0 license.
数据集概述:ovos-wake-word-bench-mlsw-negatives-nl-NL
基本信息
- 许可证:Apache-2.0
- 语言:荷兰语(nl)
数据集性质
本数据集并非音频语料库,而是为 OVOS 插件竞技场(Plugin Arena)设计的样本集清单(sample-set manifest)。它不直接提供音频文件,而是包含一组经过种子随机化的、可复现的音频片段标识符列表。
数据来源与选择方式
- 来源语料库:Multilingual Spoken Words Corpus(MLCommons 出品,许可证为 CC-BY-4.0)
- 选择机制:通过记录随机种子和源数据行数,确保每次使用该负样本集进行唤醒词插件评测时,所有插件都面对完全相同的音频片段
- 音频归属:原始音频不在此数据集中重新分发,仍保留在 MLCommons 平台,并遵循其自身的许可与署名条款
用途与价值
该清单主要用于唤醒词插件测试中的负样本集。通过保证所有插件在相同负样本上进行评分,使得不同插件之间的误接受率(false-accept rates)具有可比性,从而支持公平的横向评测。
更多信息
- 评测结果页面:https://openvoiceos.github.io/ovos-plugin-arena/
- 清单格式说明:参见 GitHub 仓库中的
docs/SPECIFICATION.md,仓库地址:https://github.com/OpenVoiceOS/ovos-plugin-arena





