voice-authenticity-dataset
收藏资源简介:
该数据集包含人类语音和合成语音的音频样本,共108个样本(60个人类语音,48个合成语音)。每个样本提供以下信息:贡献ID、数据来源(如人类或合成)、语言、提示ID、罗马化提示文本、说话人年龄范围、性别、所在区域、录制环境、所用的TTS引擎(仅合成语音)、语音ID、音频文件(16kHz采样率)、音频时长(秒)、提交时间戳。该数据集适用于语音合成、语音识别、说话人识别、语音风格分析、多语言语音研究等任务,也可用于比较人类语音与合成语音的差异。
This dataset contains audio samples of human speech and synthetic speech, with a total of 108 samples (60 human speech and 48 synthetic speech). Each sample provides the following information: contribution ID, data source (e.g., human or synthetic), language, prompt ID, romanized prompt text, speaker age range, gender, region, recording environment, TTS engine used (synthetic speech only), voice ID, audio file (16kHz sampling rate), audio duration (seconds), and submission timestamp. This dataset is suitable for tasks such as speech synthesis, speech recognition, speaker recognition, speech style analysis, multilingual speech research, and also for comparing differences between human speech and synthetic speech.
数据集概述:voice-authenticity-dataset
基本信息
- 数据集名称:voice-authenticity-dataset
- 数据集地址:https://huggingface.co/datasets/akash1702-eng/voice-authenticity-dataset
- 数据集大小:约23.1 MB(下载大小23,114,144字节,数据集总大小23,119,178字节)
数据集内容
该数据集包含语音音频样本,每条样本附带详细元数据(共14个字段),具体特征包括:
| 特征名 | 类型 | 说明 |
|---|---|---|
| contribution_id | string | 贡献样本的唯一标识 |
| source | string | 音频来源 |
| language | string | 语言 |
| prompt_id | string | 提示文本标识 |
| prompt_text_romanized | string | 提示文本(罗马化) |
| age_range | string | 说话者年龄范围 |
| gender | string | 说话者性别 |
| region | string | 地区 |
| environment | string | 录音环境 |
| tts_engine | string | 文本转语音引擎(合成样本) |
| voice_id | string | 语音标识 |
| audio | audio (采样率16kHz) | 音频数据 |
| duration_seconds | float64 | 音频时长(秒) |
| submitted_at | string | 提交时间 |
数据集划分
数据分为两个子集:
| 划分名称 | 样本数量 | 大小 |
|---|---|---|
| human(人类语音) | 60条 | 约13.98 MB(13,984,019字节) |
| synthetic(合成语音) | 48条 | 约9.14 MB(9,135,159字节) |
数据用途
该数据集可用于语音真实性检测任务,通过对比人类语音与合成语音样本,支持语音伪造检测、语音鉴别等相关模型训练与评估。





