Corpus Minangkabau
收藏资源简介:
In the development of language technologies such as machine translation, speech recognition, and others, the language corpus is very important as a source of training data. By having a corpus of Minangkabau and Indonesian languages, language technology developers can build models and systems that are more accurate and effective. Data for the corpus is collected from a variety of sources, including a number of websites and books that provide information in Minangkabau and Indonesian. Up to 520 Minangkabau and Indonesian sentences have been compiled in the data you want to publish, all of which are presented as rhymes and poetry.
在机器翻译、语音识别等语言技术的研发过程中,语言语料库(language corpus)作为训练数据的核心来源,具有不可替代的重要价值。借助米南加保语(Minangkabau)与印尼语(Indonesian)双语语料库,语言技术开发者能够搭建出精度更高、效能更优的模型与系统。该语料库的数据来源于多类渠道,涵盖了大量发布米南加保语与印尼语内容的网站及书籍。本次待发布的数据集共整理收录至多520组米南加保语-印尼语平行语句,所有文本均以韵文及诗歌的形式呈现。




