遇见数据集

open-llm-leaderboard/details_AIDC-ai-business__Marcoroni-13B

收藏
Hugging Face2023-09-18 更新2024-03-04 收录
官方服务:

资源简介:

该数据集是在Open LLM Leaderboard上对模型AIDC-ai-business/Marcoroni-13B进行评估时自动创建的。数据集由61个配置组成,每个配置对应一个评估任务。数据集由2次运行创建,每次运行在每个配置中作为一个特定的分割存在,train分割始终指向最新的结果。此外,还有一个名为results的配置存储了所有运行的聚合结果,用于计算和显示Open LLM Leaderboard上的聚合指标。文件还提供了如何加载运行中的详细信息的说明,并提供了特定运行的最新结果。

This dataset was automatically created during the evaluation of the model AIDC-ai-business/Marcoroni-13B on the Open LLM Leaderboard. It consists of 61 configurations, each corresponding to one evaluation task. The dataset is generated from 2 runs, where each run exists as a specific split under each configuration, and the `train` split always points to the most recent results. Additionally, there is a configuration named `results` that stores the aggregated results of all runs, which are used to calculate and display the aggregate metrics on the Open LLM Leaderboard. The documentation also provides instructions on how to load detailed information from the runs, as well as the most recent results for a specific run.

提供机构:
open-llm-leaderboard
原始信息汇总

数据集概述

数据集简介

该数据集是在模型 AIDC-ai-business/Marcoroni-13BOpen LLM Leaderboard 上的评估运行期间自动创建的。

数据集组成

  • 数据集包含 61 个配置,每个配置对应一个评估任务。
  • 数据集从 2 次运行中创建,每次运行可以在每个配置中找到特定的分割,分割名称使用运行的时间戳。
  • "train" 分割始终指向最新的结果。
  • 一个额外的配置 "results" 存储所有运行的聚合结果,用于计算和显示 Open LLM Leaderboard 上的聚合指标。

数据加载示例

python from datasets import load_dataset data = load_dataset("open-llm-leaderboard/details_AIDC-ai-business__Marcoroni-13B", "harness_truthfulqa_mc_0", split="train")

最新结果

以下是 2023-09-18T15:05:14.072037 运行的最新结果

python { "all": { "acc": 0.5968939242056442, "acc_stderr": 0.03397009205870784, "acc_norm": 0.6007957237246586, "acc_norm_stderr": 0.033948145854358645, "mc1": 0.4186046511627907, "mc1_stderr": 0.017270015284476855, "mc2": 0.5769635027861147, "mc2_stderr": 0.015727623906231773 }, "harness|arc:challenge|25": { "acc": 0.590443686006826, "acc_stderr": 0.014370358632472447, "acc_norm": 0.6245733788395904, "acc_norm_stderr": 0.014150631435111726 }, "harness|hellaswag|10": { "acc": 0.6366261700856403, "acc_stderr": 0.004799882248494813, "acc_norm": 0.8327026488747261, "acc_norm_stderr": 0.003724783389253322 }, "harness|hendrycksTest-abstract_algebra|5": { "acc": 0.3, "acc_stderr": 0.046056618647183814, "acc_norm": 0.3, "acc_norm_stderr": 0.046056618647183814 }, "harness|hendrycksTest-anatomy|5": { "acc": 0.5407407407407407, "acc_stderr": 0.04304979692464242, "acc_norm": 0.5407407407407407, "acc_norm_stderr": 0.04304979692464242 }, "harness|hendrycksTest-astronomy|5": { "acc": 0.5921052631578947, "acc_stderr": 0.039993097127774734, "acc_norm": 0.5921052631578947, "acc_norm_stderr": 0.039993097127774734 }, "harness|hendrycksTest-business_ethics|5": { "acc": 0.54, "acc_stderr": 0.05009082659620332, "acc_norm": 0.54, "acc_norm_stderr": 0.05009082659620332 }, "harness|hendrycksTest-clinical_knowledge|5": { "acc": 0.5962264150943396, "acc_stderr": 0.030197611600197946, "acc_norm": 0.5962264150943396, "acc_norm_stderr": 0.030197611600197946 }, "harness|hendrycksTest-college_biology|5": { "acc": 0.6597222222222222, "acc_stderr": 0.039621355734862175, "acc_norm": 0.6597222222222222, "acc_norm_stderr": 0.039621355734862175 }, "harness|hendrycksTest-college_chemistry|5": { "acc": 0.37, "acc_stderr": 0.04852365870939099, "acc_norm": 0.37, "acc_norm_stderr": 0.04852365870939099 }, "harness|hendrycksTest-college_computer_science|5": { "acc": 0.51, "acc_stderr": 0.05024183937956912, "acc_norm": 0.51, "acc_norm_stderr": 0.05024183937956912 }, "harness|hendrycksTest-college_mathematics|5": { "acc": 0.37, "acc_stderr": 0.04852365870939099, "acc_norm": 0.37, "acc_norm_stderr": 0.04852365870939099 }, "harness|hendrycksTest-college_medicine|5": { "acc": 0.6011560693641619, "acc_stderr": 0.037336266553835096, "acc_norm": 0.6011560693641619, "acc_norm_stderr": 0.037336266553835096 }, "harness|hendrycksTest-college_physics|5": { "acc": 0.3235294117647059, "acc_stderr": 0.04655010411319616, "acc_norm": 0.3235294117647059, "acc_norm_stderr": 0.04655010411319616 }, "harness|hendrycksTest-computer_security|5": { "acc": 0.71, "acc_stderr": 0.04560480215720685, "acc_norm": 0.71, "acc_norm_stderr": 0.04560480215720685 }, "harness|hendrycksTest-conceptual_physics|5": { "acc": 0.5148936170212766, "acc_stderr": 0.03267151848924777, "acc_norm": 0.5148936170212766, "acc_norm_stderr": 0.03267151848924777 }, "harness|hendrycksTest-econometrics|5": { "acc": 0.39473684210526316, "acc_stderr": 0.045981880578165414, "acc_norm": 0.39473684210526316, "acc_norm_stderr": 0.045981880578165414 }, "harness|hendrycksTest-electrical_engineering|5": { "acc": 0.5586206896551724, "acc_stderr": 0.04137931034482757, "acc_norm": 0.5586206896551724, "acc_norm_stderr": 0.04137931034482757 }, "harness|hendrycksTest-elementary_mathematics|5": { "acc": 0.36243386243386244, "acc_stderr": 0.02475747390275206, "acc_norm": 0.36243386243386244, "acc_norm_stderr": 0.02475747390275206 }, "harness|hendrycksTest-formal_logic|5": { "acc": 0.38095238095238093, "acc_stderr": 0.043435254289490965, "acc_norm": 0.38095238095238093, "acc_norm_stderr": 0.043435254289490965 }, "harness|hendrycksTest-global_facts|5": { "acc": 0.42, "acc_stderr": 0.049604496374885836, "acc_norm": 0.42, "acc_norm_stderr": 0.049604496374885836 }, "harness|hendrycksTest-high_school_biology|5": { "acc": 0.6580645161290323, "acc_stderr": 0.026985289576552742, "acc_norm": 0.6580645161290323, "acc_norm_stderr": 0.026985289576552742 }, "harness|hendrycksTest-high_school_chemistry|5": { "acc": 0.458128078817734, "acc_stderr": 0.03505630140785741, "acc_norm": 0.458128078817734, "acc_norm_stderr": 0.03505630140785741 }, "harness|hendrycksTest-high_school_computer_science|5": { "acc": 0.59, "acc_stderr": 0.04943110704237101, "acc_norm": 0.59, "acc_norm_stderr": 0.04943110704237101 }, "harness|hendrycksTest-high_school_european_history|5": { "acc": 0.7151515151515152, "acc_stderr": 0.035243908445117815, "acc_norm": 0.7151515151515152, "acc_norm_stderr": 0.035243908445117815 }, "harness|hendrycksTest-high_school_geography|5": { "acc": 0.7727272727272727, "acc_stderr": 0.029857515673386417, "acc_norm": 0.7727272727272727, "acc_norm_stderr": 0.029857515673386417 }, "harness|hendrycksTest-high_school_government_and_politics|5": { "acc": 0.8601036269430051, "acc_stderr": 0.025033870583015178, "acc_norm": 0.8601036269430051, "acc_norm_stderr": 0.025033870583015178 }, "harness|hendrycksTest-high_school_macroeconomics|5": { "acc": 0.6076923076923076, "acc_stderr": 0.02475600038213095, "acc_norm": 0.6076923076923076, "acc_norm_stderr": 0.02475600038213095 }, "harness|hendrycksTest-high_school_mathematics|5": { "acc": 0.32592592592592595, "acc_stderr": 0.028578348365473072, "acc_norm": 0.32592592

搜集汇总
数据集介绍
open-llm-leaderboard/details_AIDC-ai-business__Marcoroni-13B 数据集图片
构建方式
在大型语言模型评估领域,Open LLM Leaderboard 作为权威的基准平台,其评估过程会动态生成结构化数据集。本数据集即是在对 AIDC-ai-business/Marcoroni-13B 模型进行系统性评测的过程中自动构建而成。数据集整合了 61 个独立配置,每个配置对应一项具体的评测任务,涵盖了从常识推理到专业知识的多维度测试。构建过程历经两次独立运行,每次运行的评测结果均被保存为一个独立的数据分割,分割名称以运行时间戳精确标识,而 'train' 分割则始终指向最新一次的评测输出。此外,额外增设的 'results' 配置用于汇总所有运行的整体指标,为排行榜的最终呈现提供数据基础。
特点
该数据集的核心特征在于其精细的结构化设计与版本追溯能力。每个配置下的数据分割均按时间戳严格区分,使得研究者能够回溯模型在不同时间点的性能表现,便于进行纵向对比与分析。数据集不仅包含如 ARC-Challenge、HellaSwag 等通用基准任务的细粒度评分,还涵盖了 Hendrycks 测试中涵盖医学、法律、哲学等 57 个学科领域的专业知识评估结果,展现了模型在多元知识维度上的能力图谱。每个任务条目均附带了准确率及其标准误差等统计指标,确保了评估结果的可信度与可复现性。
使用方法
使用本数据集时,研究者可通过 Hugging Face 的 datasets 库便捷加载特定任务配置的评测细节。例如,调用 load_dataset 函数并指定配置名称如 'harness_truthfulqa_mc_0' 以及所需的分割(如 'train' 以获取最新结果),即可获取该任务下模型在每个样本上的详细得分。通过遍历不同配置,能够系统性地分析模型在各类任务上的表现差异。此外,'results' 配置提供了聚合后的总览指标,便于快速评估模型的整体性能。数据以 Parquet 格式存储,兼顾了高效存取与大数据量处理的优势。
背景与挑战
背景概述
在大语言模型(LLM)能力评估领域,Open LLM Leaderboard由Hugging Face团队于2023年创建,旨在通过标准化基准测试推动模型性能的透明比较。该数据集针对AIDC-ai-business开发的Marcoroni-13B模型,围绕其在不同自然语言理解与推理任务上的表现展开系统评估。核心研究问题在于如何量化13B参数级别模型在多项复杂基准(如ARC挑战集、HellaSwag常识推理及涵盖57个学科的大规模多任务语言理解测试)中的泛化能力。作为开源社区的重要参考,该评估结果不仅为模型开发者提供了细粒度性能反馈,更推动了LLM评估范式的规范化,对后续模型优化与基准选择产生了显著影响。
当前挑战
该数据集所应对的领域挑战在于大语言模型性能评估的碎片化与不可复现性。传统评估常因任务选择、提示格式或采样策略的差异导致结果偏差,而Open LLM Leaderboard通过固定配置(如61个评估任务、统一采样数)实现标准化比较,解决了跨模型公平性难题。构建过程中,挑战集中于多轮运行结果的结构化整合:需将不同时间戳的评估快照(如2023年9月的两轮测试)映射为独立数据集拆分,并保证最新结果始终指向“train”拆分。此外,面对57个学科子任务的庞杂指标(如acc_norm与mc2),数据格式的设计需兼顾细粒度可追溯性与聚合可读性,这对Parquet文件的版本管理与元数据一致性提出了极高要求。
常用场景
经典使用场景
在开放大语言模型评测的学术浪潮中,Open LLM Leaderboard 评测数据集为 Marcoroni-13B 这类模型提供了标准化性能评估的经典平台。该数据集涵盖 61 个评测任务配置,囊括 ARC-Challenge、HellaSwag 以及涵盖 57 个学科的 MMLU(HendrycksTest)等核心基准,同时包含 TruthfulQA 等真实性评估任务。研究者通过加载特定配置(如 harness_arc_challenge_25)并调用最新运行分片,即可精确复现模型在常识推理、知识掌握与事实一致性等维度的表现,从而系统性地衡量模型的综合语言理解能力。
实际应用
在实际应用中,该数据集的核心价值在于为模型选型与部署提供量化依据。开发者和企业可依据 Marcoroni-13B 在 MMLU 高中政府与政治(86.01%)或世界宗教(80.70%)等任务上的高准确率,判断其在教育问答、知识检索系统中的适用性。同时,TruthfulQA 的 MC1 和 MC2 指标(分别为 41.86% 和 57.70%)帮助评估模型在事实性输出上的可靠性,这对于医疗、法律等高风险场景的落地至关重要。评测结果直接指导了模型在聊天机器人、智能辅导等真实产品中的优化方向。
衍生相关工作
该数据集衍生了一系列关于大语言模型能力边界与评测方法论的经典工作。基于其多任务结构,研究者开发了针对特定学科(如医学、法学)的细粒度分析工具,并催生了诸如“MMLU 基准下的模型知识图谱”等后续研究。此外,Marcoroni-13B 的多次评测运行记录(如 2023-09-18 与 2023-09-11 的差异)为探索模型性能的时序稳定性提供了数据支撑,推动了关于评测噪声与置信区间分析的学术讨论。这些工作共同深化了学界对模型泛化能力与评测范式的理解。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务