japanese-anime-speech-v2
收藏资源简介:
japanese-anime-speech-v2是一个音频-文本数据集,旨在训练自动语音识别模型。该数据集包含300,506个音频片段及其对应的转录文本,来源于视觉小说。数据集的目标是提高自动语音识别模型(如OpenAI的Whisper)对动漫和其他类似日本媒体对话的转录准确性。音频格式为mp3,采样率为16000Hz,平均音频长度为5.5秒。这是japanese-anime-speech-v2系列的第一版,与前一版本相比,音频质量有所调整,未过滤NSFW内容。数据集主要由女性声音组成,词汇围绕爱情、关系和幻想等主题,可能不完全反映现实世界的说话模式。未来计划包括创建安全工作和NSFW内容的分离,改进文本格式,以及扩展数据集来源。
japanese-anime-speech-v2 is an audio-text dataset intended for training automatic speech recognition (ASR) models. It contains 300,506 audio clips and their corresponding transcriptions, sourced from visual novels. The dataset aims to enhance the transcription accuracy of automatic speech recognition models such as OpenAI's Whisper for dialogues in anime and other similar Japanese media. The audio is in MP3 format with a sampling rate of 16000 Hz, and the average length of each audio clip is 5.5 seconds. This is the first iteration of the japanese-anime-speech-v2 series; compared to the prior version, the audio quality has been adjusted, and no NSFW content has been filtered. The dataset predominantly features female voices, with its vocabulary centered on themes including love, relationships, and fantasy, and may not fully reflect real-world speech patterns. Future plans for the dataset include creating a separation between safe-for-work and NSFW content, improving text formatting, and expanding the dataset's source materials.
Japanese Anime Speech Dataset V2
概述
japanese-anime-speech-v2 是一个用于训练自动语音识别模型的音频-文本数据集。该数据集包含 292,637 个音频片段 及其对应的转录文本,来源于各种视觉小说。
数据集信息
- 音频-文本对数目: 292,637
- 安全内容音频时长: 397.54小时 (86.8%)
- 非安全内容音频时长: 52.36小时 (13.2%)
- 平均安全内容音频长度: 5.3秒
- 数据来源: 视觉小说
- 音频格式: mp3 (128kbps)
- 最新版本: V2 - 2024年6月29日
数据集特点
- 音频特征:
- 采样率: 16000 Hz
- 文本特征:
- 数据类型: 字符串
数据集分割
- 安全内容 (sfw):
- 字节数: 19174765803.112
- 样本数: 271788
- 非安全内容 (nsfw):
- 字节数: 2864808426.209
- 样本数: 20849
数据集大小
- 下载大小: 24379492733 字节
- 数据集大小: 22039574229.321 字节
配置
- 默认配置:
- 安全内容文件路径: data/sfw-*
- 非安全内容文件路径: data/nsfw-*
版本变更
- 从 V1 到 V2 的变化:
- 数据集大小显著增加,从 73,004 增加到 292,637 个音频-文本对
- 音频格式从 mp3 (192kbps) 改为 mp3 (128kbps),以提高存储效率
- 安全内容和非安全内容分为不同的分割
- 重复字符已规范化
- 删除了不含对话的音频行
- 删除了低质量的音频行
偏差与限制
- 数据集主要来源于视觉小说,导致性别偏向女性声音,且词汇围绕爱情、关系和幻想等主题
- 音频质量较高,可能导致与现实世界说话模式不完全一致
- 包含非安全内容,不适用于所有应用场景
- 转录文本未进行格式化或清理,可能影响部分文本样本的质量
未来计划
- 继续扩展数据集,包括更多来源
使用与引用
- 数据集对商业和非商业用途开放
- 使用时无需强制引用,但建议在衍生作品中提供超链接




