遇见数据集

Okey Aura Wake-up Word Dataset

收藏
Zenodo2024-05-20 更新2026-05-26 收录
官方服务:

资源简介:

Speech dataset for wake-up word (WuW) detection in Telefónica's home assistant, Aura. It contains 1247 utterances (1.4 hours) from ~80 speakers. Speakers pronounce the wake-up word itself "Okey Aura", plus other sentences that might be similar, or not, to "Okey Aura". This dataset contains rich metadata annotations, so it is possible to study diverse factors and biases that might affect wake-up word detection performance: accent, gender, prosody/emotion, room size, distance to the microphone, etc. Besides, it also contains recordings of sentences that are phonetically similar to "Okey Aura", like "Porque Laura..." or "... como Aura...", to experiment with difficult sentences.

用于西班牙电信(Telefónica)旗下家用助手Aura的唤醒词(Wake-up Word,WuW)检测任务的语音数据集。该数据集包含来自约80位说话者的1247条语音样本,总时长1.4小时。说话者不仅会朗读唤醒词本身"Okey Aura",还会朗读与"Okey Aura"语音相似或不相似的其他语句。 本数据集附带丰富的元数据标注,可用于研究各类可能影响唤醒词检测性能的因素与偏差,涵盖口音、性别、韵律/情感、房间尺寸、与麦克风的间距等。此外,数据集还收录了语音层面与"Okey Aura"相似的语句,例如"Porque Laura..."或"... como Aura...",用于开展高难度语句场景下的实验。

提供机构:
Zenodo
创建时间:
2024-05-20
二维码
社区交流群
二维码
科研交流群
商业服务