ovos-wake-word-bench-mlsw-negatives-pt-PT
收藏资源简介:
该数据集是OVOS唤醒词插件竞技场(OVOS Plugin Arena)的一个负样本清单,而非音频语料库。它包含从Multilingual Spoken Words Corpus(MLCommons, CC-BY-4.0)中经种子随机选取的剪辑标识符列表,并记录了种子和源行数,以确保不同唤醒词插件在完全相同的一组剪辑上进行评分,从而使假接受率具有可比性。音频文件未重新分发,仍保留在原始数据集中。该清单以Apache-2.0许可证发布,与竞技场其他基准测试仓库一致。数据集语言为葡萄牙语(欧洲葡萄牙语)。
This dataset is a negative sample list for the OVOS Wake Word Plugin Arena, not an audio corpus. It contains clip identifiers randomly selected from the Multilingual Spoken Words Corpus (MLCommons, CC-BY-4.0) with seeds and source row counts, ensuring that different wake word plugins are scored on the exact same set of clips for comparable false acceptance rates. Audio files are not redistributed and remain in the original dataset. The list is released under the Apache-2.0 license, consistent with other benchmark repositories in the arena. The dataset language is Portuguese (European).
数据集概述
该数据集为 ovos-wake-word-bench-mlsw-negatives-pt-PT,由 OpenVoiceOS 发布,遵循 Apache-2.0 许可证,语言为葡萄牙语(pt)。
核心性质
- 这是一个 样本清单(manifest),并非音频语料库本身,用于 OVOS 插件竞技场(Plugin Arena)的基准测试。
- 它包含从 多语言口语词语库(Multilingual Spoken Words Corpus,MLCommons,CC-BY-4.0) 中选出的、可复现的片段标识符列表,并记录了随机种子和源数据行数。
用途与技术细节
- 旨在确保所有唤醒词插件在评估时使用完全相同的负面样本片段,从而使得不同插件之间的误接受率(false-accept rate)具有可比性。
- 音频内容不在此数据集中重新分发,仍托管于 MLCommons 平台,并遵循其自身的许可和署名条款。
附加信息
- 测试结果页面:https://openvoiceos.github.io/ovos-plugin-arena/
- 清单格式规范:见 https://github.com/OpenVoiceOS/ovos-plugin-arena 中的
docs/SPECIFICATION.md文件。





