Mixed Cantonese and English (MCE) audio dataset
收藏资源简介:
MCE数据集是由香港科技大学和南京航空航天大学合作创建的,专注于混合粤语和英语的自动语音识别研究。该数据集包含34.8小时的音频文件,涵盖日常生活中的18个主题,由GPT-4生成文本信息,并由志愿者根据这些文本录制音频。MCE数据集的创建旨在解决现有数据集中缺乏混合语言表达的问题,特别是在粤语和英语混合使用频繁的香港地区。该数据集的应用领域包括提升自动语音识别系统在处理混合语言环境中的性能,尤其是在粤语和英语的混合使用场景中。
The MCE Dataset was co-developed by The Hong Kong University of Science and Technology and Nanjing University of Aeronautics and Astronautics, focusing on automatic speech recognition (ASR) research for code-switched Cantonese and English. This dataset contains 34.8 hours of audio recordings spanning 18 daily life topics. Its textual content was generated by GPT-4, and audio data was collected via volunteer recordings based on these generated texts. The MCE Dataset was created to address the lack of code-switched language datasets, especially in regions like Hong Kong where Cantonese and English are frequently used in mixed contexts. Its applications include improving the performance of automatic speech recognition systems in code-switched linguistic environments, particularly in scenarios involving frequent mixed use of Cantonese and English.




