MM-Eval
收藏资源简介:
MM-Eval数据集由中国民族大学开发,专门用于评估大型语言模型(LLMs)在现代蒙古语中的表现。该数据集包含1840条数据,分为四个层次:语法、语义、知识和推理。数据集的内容主要来源于《现代蒙古语教材I》,并结合了WebQSP和MGSM数据集进行丰富。数据集的创建过程包括从教材中提取句子,进行数据清洗和手动校正,以及使用ChatGPT API生成和验证数据。MM-Eval数据集的应用领域主要集中在低资源语言的NLP研究和LLMs的性能评估,旨在解决当前模型在处理蒙古语时的不足。
MM-Eval dataset was developed by Minzu University of China, specifically tailored for evaluating the performance of Large Language Models (LLMs) in Modern Mongolian. This dataset comprises 1,840 instances classified into four dimensions: grammar, semantics, knowledge, and reasoning. Its content is primarily derived from *Modern Mongolian Textbook I*, and augmented by integrating the WebQSP and MGSM datasets. The development pipeline of the MM-Eval dataset includes extracting sentences from the textbook, conducting data cleaning and manual correction, as well as generating and validating data via the ChatGPT API. The MM-Eval dataset is mainly applied in NLP research for low-resource languages and LLMs performance evaluation, aiming to address the current limitations of existing models when processing Mongolian.




