遇见数据集

蒙古语喀尔喀方言语音数据集

收藏
官方服务:

资源简介:

蒙古语喀尔喀方言语料库是一个包含大量蒙古语喀尔喀方言的录音和对应的文本标注的语音语料库,项目组首先搜集了涵盖新闻、教育、旅游、日常用语等多个领域的西里尔蒙古文文本语料,经过文本单词长度筛选以及人工校对文字后,整理了共计5万句西里尔蒙古文语句,作为喀尔喀方言蒙古语语音库录制的文本语料。国外合作单位蒙古国科学院数学与数字技术研究所雇用了1100多名蒙古国各地区的说话人参与喀尔喀方言蒙古语语音录制。此数据集可用于各种语言学和人文学科研究,如语音识别、语音合成、民族学、人类学、历史学等领域。数据量为300小时,录制时间为2020年5月到9月。

The Khalkha Mongolian dialect corpus is a speech corpus comprising a vast collection of Khalkha Mongolian dialect audio recordings and their corresponding manually annotated text transcripts. Initially, the project team collected Cyrillic Mongolian text corpora spanning multiple domains including news, education, tourism, daily conversations and others. After filtering based on word length and conducting manual proofreading of the texts, a total of 50,000 Cyrillic Mongolian sentences were compiled as the reference text materials for recording the Khalkha Mongolian dialect speech corpus. The foreign cooperating institution, the Institute of Mathematics and Digital Technology of the Mongolian Academy of Sciences, recruited over 1,100 speakers from various regions of Mongolia to participate in the speech recording work. This dataset can be applied to various linguistic and humanities research fields, such as speech recognition, speech synthesis, ethnology, anthropology, history and more. The total duration of the dataset is 300 hours, and the recording was carried out from May to September 2020.

提供机构:
内蒙古大学
搜集汇总
数据集介绍
蒙古语喀尔喀方言语音数据集 数据集图片
背景与挑战
背景概述
该数据集是一个蒙古语喀尔喀方言的语音语料库,包含300小时的录音和对应的文本标注,基于5万句西里尔蒙古文语料,由1100多名蒙古国说话人于2020年5月至9月录制。它适用于语音识别、语音合成等语言学和人文学科研究。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务