agarwalayushi/hinglish
收藏资源简介:
Hinglish Concatenated Audio Dataset是一个大规模、经过清理和注释的语音数据集,涵盖印地语、Hinglish(印地语-英语代码转换)和印度英语。该数据集由14个公共语料库和原始自定义录音编译而成,统一为一个具有一致模式的Parquet数据集。数据集包含815,171个音频片段,总计超过2,264小时的录音,来自6,304个独特的说话者。音频格式为嵌入Parquet的WAV文件,支持ASR、TTS微调、语音克隆和语音研究等任务。数据集经过重新分段以去除静音和交叉对话,转录文本已标准化为Unicode NFC,并添加了语言标签(如`<hi-en>`表示代码转换的语句)。
A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema. The dataset contains 815,171 audio clips totaling over 2,264 hours of recordings from 6,304 unique speakers. Audio is stored as WAV embedded in Parquet, supporting tasks like ASR, TTS fine-tuning, voice cloning, and speech research. The dataset has been re-segmented to remove silence and cross-talk, with transcripts normalized to Unicode NFC and language tags added (e.g., `<hi-en>` for code-switched utterances).




