遇见数据集

llmeval-fdu/LLMEval-Fair

收藏
Hugging Face2026-05-09 更新2026-05-31 收录
官方服务:

资源简介:

LLMEval-Fair是一个大规模纵向研究数据集,专注于大型语言模型(LLM)评估的鲁棒性和公平性。它基于一个专有的题库,包含超过22万个研究生级别的中文问题,涵盖13个学术学科。该数据集是原始题库的公开子集,旨在解决现有静态基准测试的四个关键问题:抗污染性(通过动态采样未见的测试集)、防作弊架构(包括单问题序列分发、部分封闭题库和跨机构重复检测)、校准的LLM作为评判者(与人类专家约90%的一致性)以及纵向覆盖(在30个月内测试了近60个领先模型,如GPT-5、Claude-Sonnet-4.5等)。数据集包括经济科学、工程、历史、法律、文学、管理科学、医学、军事科学、自然科学、哲学科学、教育等学科,每个学科下按问题类型(如多项选择、简答、判断对错等)组织数据。

LLMEval-Fair is a large-scale longitudinal study on the robustness and fairness of LLM evaluation, built on a proprietary bank of 220,000+ graduate-level Chinese questions spanning 13 academic disciplines. This is the publicly released subset of that bank, addressing four key issues with existing static benchmarks: contamination resistance (dynamically samples unseen test sets), anti-cheating architecture (single-question serial dispatch, partially closed item bank, cross-institutional duplicate detection), calibrated LLM-as-a-judge (~90% agreement with human experts), and longitudinal coverage (nearly 60 leading models tested over 30 months, including GPT-5, Claude-Sonnet-4.5, etc.). The dataset covers disciplines such as Economic Sciences, Engineering, History, Law, Literature, Management Sciences, Medicine, Military Science, Natural Sciences, Philosophical Sciences, and Education, with data organized by question type (e.g., Multiple Choice, Short Answer, True-False).

提供机构:
llmeval-fdu
二维码
社区交流群
二维码
科研交流群
商业服务