ovos-wake-word-bench-mlsw-negatives-fr-FR
收藏资源简介:
该数据集名为 ovos-wake-word-bench-mlsw-negatives-fr-FR,是一个用于 OVOS 插件竞技场(OVOS Plugin Arena)的样本集清单(manifest),而非音频语料库。它包含从 Multilingual Spoken Words Corpus(MLCommons 提供,采用 CC-BY-4.0 许可证)中选取的剪辑标识符列表,并附有种子和源行计数,以确保可重复性。通过该清单,所有针对这些负样本评分的唤醒词插件均在同一组剪辑上进行评估,从而使得假接受率在不同插件之间具有可比性。音频文件本身并未重新分发,仍保留在 MLCommons 下,需遵守其原始许可和归属条款。本清单采用 Apache-2.0 许可证,与竞技场其他基准仓库一致。
This dataset is named ovos-wake-word-bench-mlsw-negatives-fr-FR, a manifest of sample IDs for the OVOS Plugin Arena, not an audio corpus. It contains a list of clip identifiers selected from the Multilingual Spoken Words Corpus (provided by MLCommons under CC-BY-4.0 license), along with seed and source row counts to ensure reproducibility. Using this manifest, all wake-word plugins scored against these negatives are evaluated on the same set of clips, making false acceptance rates comparable across plugins. The audio files themselves are not redistributed and remain under MLCommons with their original license and attribution. This manifest is licensed under Apache-2.0, consistent with other benchmark repositories in the arena.
数据集概述
数据集名称:ovos-wake-word-bench-mlsw-negatives-fr-FR
许可协议:Apache-2.0
语言:法语(fr)
核心说明
- 该数据集是 OVOS Plugin Arena 的 样本清单(manifest),并非音频语料库本身。
- 它包含一个 可复现的、基于种子(seed)的片段标识符列表,这些片段选自 Multilingual Spoken Words Corpus(MLCommons,许可协议为 CC-BY-4.0)。
- 清单中记录了 种子值 和 源数据行数,确保所有参与测试的唤醒词插件都使用完全相同的音频片段进行评分,从而使 误唤醒率(false-accept rates) 具有跨插件可比性。
数据与版权说明
- 音频文件未被重新分发,仍保留在 MLCommons 原始数据集中,受其自身许可和署名条款约束。
- 该清单本身以 Apache-2.0 许可发布,与 Arena 其他基准测试仓库保持一致。
相关链接
- 测试结果:https://openvoiceos.github.io/ovos-plugin-arena/
- 清单格式规范:
docs/SPECIFICATION.md,位于 GitHub 仓库 https://github.com/OpenVoiceOS/ovos-plugin-arena




