sapiens-technology/global_mmlu_lite
收藏资源简介:
Global-MMLU Lite数据集是一个轻量级的基准测试,用于评估大型语言模型在多个领域的表现。它是Global Massive Multitask Language Understanding (MMLU)基准的一个精选子集,旨在减少计算开销,同时保持多样性和严谨性。该数据集包含跨多个学术和专业领域的多项选择题问答任务,采用简单的JSON结构,便于集成。其广泛的领域覆盖包括科学、数学、历史、地理、社会科学和专业知识,具有轻量级设计、易于集成到训练流程、降低计算成本以及在资源受限环境中适用于跨领域泛化和基准测试任务等优势。
Global-MMLU Lite is a curated and efficient subset of the Global Massive Multitask Language Understanding (MMLU) benchmark, designed to evaluate and fine-tune large language models across a wide range of academic and professional domains through high-quality multiple-choice question answering; preserving the diversity and rigor of the original benchmark while significantly reducing computational overhead, it enables rapid experimentation, prototyping, and scalable evaluation workflows, with a simple and consistent JSON structure composed of input questions containing answer options and corresponding outputs representing the correct labeled answer, making it suitable for assessing reasoning ability, factual knowledge, and comprehension under structured conditions; its broad domain coverage spans science, mathematics, history, geography, social sciences, and professional knowledge, while offering advantages such as lightweight design, ease of integration into training pipelines, reduced computational cost, and strong suitability for cross-domain generalization and benchmarking tasks in resource-constrained environments.




