Shabarigirish/codeswitch-pairs-lase-indian
收藏资源简介:
Codeswitch Pairs LASE — 印度口音保留语料库是一个包含1369个跨脚本话语对的数据集,专为研究口音条件性现象而设计。它使用8个ElevenLabs印度英语多语言语音合成,展示了现成编码器在处理印度口音语音时的特性:无论脚本如何,这些编码器会将印度口音语音紧密聚类,而西方口音则表现出较大的脚本条件性差异。每个数据行代表一个合成的话语及其元数据,通过voice_id字段在评估时重建跨脚本对(同一语音、不同脚本)。数据集包含英语(en)、印地语(hi)、泰卢固语(te)和泰米尔语(ta)四种语言,音频为16 kHz单声道WAV格式,每段约2秒。质量门槛基于WavLM余弦相似度≥0.90。数据通过ElevenLabs Multilingual v2 API合成,语音来自公共目录,文本为短通用英语短语的翻译或转写。该数据集遵循CC-BY-4.0许可证,并支持音频分类和文本到语音等任务。
Codeswitch Pairs LASE — Indian-accent held-out corpus is a dataset containing 1369 cross-script utterance pairs, designed to study accent-conditional phenomena. It is synthesized using 8 ElevenLabs Indian-English Multilingual voices, highlighting the finding that off-the-shelf encoders cluster Indian-accent voices closely regardless of script, while Western voices exhibit large script-conditional gaps. Each row represents one synthesized utterance with metadata, and pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script). The dataset includes four languages: English (en), Hindi (hi), Telugu (te), and Tamil (ta). Audio is in 16 kHz mono WAV format, approximately 2 seconds per utterance. A quality gate is applied with WavLM-cosine similarity ≥ 0.90. Data is synthesized via the ElevenLabs Multilingual v2 API, using public voice IDs and short generic English phrases translated/transliterated into target scripts. It is licensed under CC-BY-4.0 and supports tasks such as audio-classification and text-to-speech.




