ovos-wake-word-bench-mlsw-negatives-pl-PL
收藏资源简介:
该数据集是一个样本集清单,专为OVOS插件竞技场设计,用于评估唤醒词检测插件的负样本性能。它包含从Multilingual Spoken Words Corpus(MLCommons, CC-BY-4.0)中选取的波兰语(pl-PL)剪辑标识符列表,该列表基于预设种子随机生成且可复现,并记录了种子和源数据行数,确保不同插件在完全相同的剪辑上进行评分,从而实现假接受率(false-accept rates)的公平比较。音频文件本身不在此仓库中重新分发,需从原始来源获取。本数据集采用Apache-2.0许可证发布。
This dataset is a sample list designed for the OVOS Plugin Arena to evaluate the negative sample performance of wake word detection plugins. It contains a list of Polish (pl-PL) clip identifiers selected from the Multilingual Spoken Words Corpus (MLCommons, CC-BY-4.0), generated randomly based on a preset seed and reproducible, with seed and source data row counts recorded. This ensures that different plugins are scored on exactly the same clips, enabling fair comparison of false-accept rates. The audio files themselves are not redistributed in this repository and must be obtained from the original source. This dataset is released under the Apache-2.0 license.
数据集概述
该数据集是一个用于 OVOS 插件竞技场(Plugin Arena)的样本清单(manifest),而非实际的音频语料库。
核心内容
- 用途:为 OVOS 唤醒词插件提供一组固定的负样本(negatives)测试集,用于评估各插件的误唤醒率(false-accept rates)。
- 数据来源:从 Multilingual Spoken Words Corpus(MLCommons,CC-BY-4.0 许可)中选取的片段标识符列表。
- 选择机制:记录种子值(seed)和源行数,确保可复现;每个唤醒词插件在完全相同的音频片段上进行评分,从而实现公平比较。
- 音频状态:音频文件不在此重新分发,仍归 MLCommons 所有,并遵循其原始许可与署名要求。
技术属性
- 语言:波兰语(pl)
- 许可证:Apache-2.0(与竞技场其他基准仓库一致)
外部链接
- 评测结果:https://openvoiceos.github.io/ovos-plugin-arena/
- 格式规范:定义于 https://github.com/OpenVoiceOS/ovos-plugin-arena 中的
docs/SPECIFICATION.md




