ovos-wake-word-bench-mlsw-negatives-it-IT
收藏资源简介:
该数据集是OVOS唤醒词基准测试(OVOS Plugin Arena)的负样本集合清单(manifest),专门用于评估意大利语唤醒词插件。它并非音频语料库,而是一个可重现的clip标识符列表,这些标识符选自MLCommons的多语言口语词汇语料库(Multilingual Spoken Words Corpus,CC-BY-4.0许可)。清单中记录了用于选择的随机种子和源数据集行数,确保所有插件在完全相同的负样本上进行评分,从而使得不同插件间的误接受率(false-accept rate)具有可比性。音频文件本身不包含在此数据集中,用户需从原始语料库中获取。该数据集适用于唤醒词检测系统的性能评估,特别是针对误接受率的标准化测试。
This dataset is a manifest of negative samples for the OVOS wake word benchmark (OVOS Plugin Arena), specifically designed to evaluate Italian wake word plugins. It is not an audio corpus, but a reproducible list of clip identifiers selected from the MLCommons Multilingual Spoken Words Corpus (CC-BY-4.0 license). The manifest records the random seed used for selection and the number of source dataset rows, ensuring that all plugins are scored on exactly the same negative samples, making the false-accept rates comparable across different plugins. The audio files themselves are not included in this dataset; users need to obtain them from the original corpus. This dataset is suitable for performance evaluation of wake word detection systems, particularly for standardized testing of false-accept rates.
ovos-wake-word-bench-mlsw-negatives-it-IT 数据集总结
基本信息
- 许可证: Apache-2.0
- 语言: 意大利语(it)
- 数据集提供方: OpenVoiceOS
数据集本质
该数据集并非音频语料库,而是一个用于 OVOS Plugin Arena 的样本集清单(manifest)。它存储了一系列可复现的、由固定随机种子生成的剪辑标识符(clip identifiers)列表。
数据来源
- 选取自 Multilingual Spoken Words Corpus(MLCommons,许可证为 CC-BY-4.0)
- 数据集中记录了随机种子和源数据行数,确保每个唤醒词插件在测试时使用完全相同的音频剪辑,从而保证不同插件之间的误接受率(false-accept rates)具有可比性
音频说明
- 音频文件并未在此数据集中重新分发,其版权归 MLCommons 所有,需遵循其自身的许可证和归属条款
用途
- 用于基准测试(benchmark),作为唤醒词插件的负样本(negatives),在 OVOS Plugin Arena 中参与评分对比
相关资源
- 基准测试结果页: https://openvoiceos.github.io/ovos-plugin-arena/
- Manifest 格式规范文档: 位于 https://github.com/OpenVoiceOS/ovos-plugin-arena 仓库中的
docs/SPECIFICATION.md




