MMLU-CF
收藏资源简介:
MMLU-CF是由微软研究院创建的一个无污染的多任务语言理解基准数据集,旨在评估大型语言模型(LLMs)的世界知识理解能力。该数据集包含20000个问题,分为10000个测试集和10000个验证集,涵盖14个领域,从2000亿份公开文档中筛选而来。数据集的创建过程包括多选题收集、清洗、难度采样、LLMs检查和无污染处理,确保数据集的多样性和高质量。MMLU-CF的应用领域主要是评估和提升LLMs在多任务环境下的表现,旨在解决现有基准数据集可能存在的数据泄露问题,提供一个更严格和可靠的评估标准。
MMLU-CF is a pollution-free multi-task language understanding benchmark dataset developed by Microsoft Research, designed to evaluate the world knowledge comprehension capabilities of large language models (LLMs). This dataset contains 20,000 questions, split into a 10,000-question test set and a 10,000-question validation set, covering 14 distinct domains, and is curated from 200 billion publicly available documents. The dataset construction pipeline encompasses multiple-choice question collection, data cleaning, difficulty sampling, LLM-based verification and pollution-free processing, which ensures the dataset's diversity and high quality. The primary applications of MMLU-CF lie in evaluating and enhancing the performance of LLMs in multi-task settings, with the goal of addressing potential data leakage issues in existing benchmark datasets and providing a more rigorous and reliable evaluation standard.




