Praxel/codeswitch-pairs-lase-indian
收藏资源简介:
Codeswitch Pairs LASE — Indian-accent held-out corpus数据集包含1369个跨脚本的语音对,由8个ElevenLabs印度-英语多语言语音合成。该数据集用于研究口音条件性发现,如现成的编码器会将印度口音的语音紧密聚类,而不考虑脚本,而西方语音则显示出较大的脚本条件性差距。数据集中的每条记录都是一个合成的语音片段及其元数据,语音对在评估时通过voice_id进行重建。支持的语言包括英语(en)、印地语(hi)、泰卢固语(te)和泰米尔语(ta)。音频格式为16 kHz单声道WAV,每个语音片段约2秒。数据集还提供了质量门限,要求WavLM-cosine ≥ 0.90与语音的参考片段相比。数据集的合成使用了ElevenLabs Multilingual v2 API,并遵循其研究/评估用途的使用条款。
The Codeswitch Pairs LASE — Indian-accent held-out corpus contains 1369 held-out cross-script utterance pairs from 8 ElevenLabs Indian-English Multilingual voices. It surfaces the accent-conditional finding: off-the-shelf encoders cluster Indian-accent voices closely regardless of script, while Western voices show large script-conditional gaps. Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id. Supported languages include English (en), Hindi (hi), Telugu (te), and Tamil (ta). Audio is 16 kHz mono WAV, ~2 s/utterance. The dataset includes a quality gate: WavLM-cosine ≥ 0.90 vs the voices reference clip. The dataset was synthesized using the ElevenLabs Multilingual v2 API and is used under their TOS for research/evaluation purposes.




