JA_Emilia_Yodas_ScribeEvents
收藏资源简介:
JA Emilia Yodas - Scribe Events (Filtered) 是一个日语语音数据集,是从 MrDragonFox/JA_Emilia_Yodas_266h 数据集中过滤得到的子集,仅包含 ElevenLabs Scribe v1 音频事件的样本。数据集主要应用于自动语音识别任务,特别关注声音爆发(如笑声、叹息等)、背景音(如背景噪音、音乐等)和其他事件(如暂停、无法识别的声音等)的标注。数据集经过过滤,保留了 4433 行数据,并对事件标注格式进行了统一处理(将 `(event)` 替换为 `[event]`)。数据集的许可证为 CC BY 4.0,适用于语音识别和音频事件检测等任务。
JA Emilia Yodas - Scribe Events (Filtered) is a Japanese speech dataset. It is a filtered subset sourced from the MrDragonFox/JA_Emilia_Yodas_266h dataset, and only includes samples corresponding to ElevenLabs Scribe v1 audio events. This dataset is primarily utilized for automatic speech recognition tasks, with particular emphasis on annotations for sound bursts (e.g., laughter, sighs), background sounds (e.g., ambient noise, music), and other events such as pauses and unrecognizable sounds. Following filtering, the dataset retains 4433 rows of data, and the event annotation format has been uniformly standardized by replacing `(event)` with `[event]`. The dataset is licensed under CC BY 4.0, and is applicable to tasks including speech recognition and audio event detection.
JA Emilia Yodas - Scribe Events (Filtered) 数据集概述
基本信息
- 数据集名称:JA Emilia Yodas - Scribe Events (Filtered)
- 许可证:CC BY 4.0
- 语言:日语 (ja)
- 任务类别:自动语音识别
- 标签:vocal-bursts, scribe-events, emilia
数据集描述
该数据集是 MrDragonFox/JA_Emilia_Yodas_266h 的一个过滤子集,仅包含具有 ElevenLabs Scribe v1 音频事件 的样本。
与源数据集的差异
- 过滤:仅保留
events_scribe字段非空的行(共保留 4433 行)。 - 括号格式统一:将
text_scribe中的(event)替换为[event]。
事件类型
- 发声事件:
<laughs>,<sighs>,<clears throat>等。 - 背景事件:
<background noise>,<music>等。 - 其他事件:
<pause>,<unintelligible>,<bleep>等。
数据来源
派生自 MrDragonFox/JA_Emilia_Yodas_266h(CC BY 4.0 许可证)。




