全球多口音英语高质量语音数据集
收藏资源简介:
该数据集融合英语播客音频与专业团队定向采集的65国口音英语语音数据,总时长达124万小时,数据规模为125 TB。收录超过42000种音色,覆盖标准英语、地域方言、非母语口音等多样化语音形态,支撑语音识别、语言研究、跨文化交流等领域前沿研究,同时该数据集近三年累计服务44个全球知名企业,覆盖商贸流通、智能驾驶、智慧金融、教育科研等领域多个场景。
This dataset integrates English podcast audio and English speech data of 65 national English accents collected by a professional team targeting these accents. It has a total duration of 1.24 million hours and a total data volume of 125 TB, containing over 42,000 unique voice timbres. The dataset covers diverse speech forms including standard English, regional dialects, non-native accents and more, supporting cutting-edge research in fields such as speech recognition, linguistic studies, cross-cultural communication and other related areas. Additionally, over the past three years, this dataset has served 44 globally renowned enterprises, covering multiple application scenarios across sectors including commercial trade circulation, intelligent driving, smart finance, education and research.




