ovos-wake-word-bench-mlsw-negatives-es-ES
收藏资源简介:
该数据集是一个样本集清单,用于OVOS插件竞技场(OVOS Plugin Arena),而非音频语料库。它包含一个可重复的、经过种子设定的剪辑标识符列表,这些标识符选自Multilingual Spoken Words Corpus(MLCommons,CC-BY-4.0),并记录了种子和源行数,以确保所有针对这些负样本进行评分的唤醒词插件都在完全相同的剪辑上评分,从而使误接受率在不同插件之间具有可比性。音频未在此重新分发,仍保留在MLCommons下,受其自身许可和署名条款约束。此清单以Apache-2.0许可证发布,与竞技场的其他基准测试仓库一致。
This dataset is a sample set list for the OVOS Plugin Arena, not an audio corpus. It contains a reproducible, seeded list of clip identifiers selected from the Multilingual Spoken Words Corpus (MLCommons, CC-BY-4.0), along with the seed and source line numbers, to ensure that all wake word plugins scoring on these negative samples are evaluated on exactly the same clips, making false acceptance rates comparable across plugins. The audio is not redistributed here and remains under MLCommons, subject to its own license and attribution terms. This list is released under the Apache-2.0 license, consistent with other benchmark repositories in the arena.
数据集概述:ovos-wake-word-bench-mlsw-negatives-es-ES
数据集类型与用途
这是一个样本集清单(manifest),而非音频语料库,专为 OVOS 插件竞技场(Plugin Arena)设计,用于评估西班牙语(es)唤醒词插件的误唤醒率(false-accept rates)。
数据来源与筛选机制
- 源数据:来自 Multilingual Spoken Words Corpus(MLCommons 发布,许可协议为 CC-BY-4.0)。
- 生成方式:通过固定随机种子,从源语料库中可重复地选取一批音频片段标识符,并记录种子值及源数据行数,确保不同唤醒词插件在完全相同的片段上接受评分,从而实现跨插件的误唤醒率可比性。
音频与授权说明
- 音频文件不在此数据集中重新分发,仍归属于 MLCommons 原始许可与署名条款下。
- 本清单本身以 Apache-2.0 许可发布,与竞技场其他基准测试仓库一致。
相关资源链接
- 评测结果:https://openvoiceos.github.io/ovos-plugin-arena/
- 清单格式规范:详见 https://github.com/OpenVoiceOS/ovos-plugin-arena 中的
docs/SPECIFICATION.md文档。
语言与授权信息
- 语言:西班牙语(es)
- 许可证:Apache-2.0




