Multitask National Speech Corpus (MNSC)
收藏资源简介:
Multitask National Speech Corpus (MNSC) 是由新加坡科技研究局信息通信研究院开发的一个大规模口语新加坡英语语料库,旨在支持多种任务,如自动语音识别、口语问答、口语对话摘要和副语言问答。该数据集包含约10,000小时的录音,涵盖了新加坡英语的多种口音和代码切换模式。数据集的创建过程包括元数据提取、现有语料库的清理以及使用大语言模型进行合成。所有测试集都经过人工注释以确保高质量和可靠性。该数据集的应用领域主要集中在多语言和代码切换的自然语言处理研究,旨在解决新加坡英语在语音技术中的独特挑战,如多语言特性、多样化的口音和复杂的句法结构。
Multitask National Speech Corpus (MNSC) is a large-scale spoken Singaporean English corpus developed by the Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (A*STAR) of Singapore. It is designed to support a variety of tasks including Automatic Speech Recognition (ASR), spoken question answering, spoken dialogue summarization, and paralinguistic question answering. This corpus contains approximately 10,000 hours of recordings, covering diverse accents and code-switching patterns of Singaporean English. The construction process of the dataset involves metadata extraction, cleaning of existing corpora, and synthesis using Large Language Models (LLMs). All test sets have undergone manual annotation to ensure high quality and reliability. The application fields of this dataset mainly focus on multilingual and code-switching natural language processing (NLP) research, aiming to tackle the unique challenges of Singaporean English in speech technology, such as its multilingual features, diverse accents, and complex syntactic structures.




