遇见数据集

open-llm-leaderboard/details_edor__Hermes-Platypus2-mini-7B

收藏
Hugging Face2023-08-27 更新2024-03-04 收录
官方服务:

资源简介:

该数据集是在Open LLM Leaderboard上对模型edor/Hermes-Platypus2-mini-7B进行评估时自动创建的。数据集包含61个配置,每个配置对应一个评估任务。数据集由1次运行创建,每次运行的结果作为特定配置中的一个分割,分割名称使用运行的时间戳。train分割始终指向最新的结果。此外,还有一个名为results的配置存储了所有运行的聚合结果,用于计算和显示Open LLM Leaderboard上的聚合指标。

This dataset was automatically created during the evaluation of the model `edor/Hermes-Platypus2-mini-7B` on the Open LLM Leaderboard. It consists of 61 configurations, each corresponding to a single evaluation task. The dataset was generated from a single run, where the results of each run are treated as a split under the corresponding configuration, and the split names are based on the timestamp of the run. The `train` split always points to the most recent evaluation results. Additionally, there is a configuration named `results` that stores the aggregated results across all runs, which is utilized to calculate and display the aggregate metrics on the Open LLM Leaderboard.

提供机构:
open-llm-leaderboard
原始信息汇总

数据集概述

数据集摘要

该数据集是在评估模型 edor/Hermes-Platypus2-mini-7BOpen LLM Leaderboard 上的评估运行期间自动创建的。数据集包含 61 个配置,每个配置对应一个评估任务。数据集从 1 次运行中创建,每个运行可以在每个配置中找到特定的分割,分割名称使用运行的时间戳。"train" 分割始终指向最新的结果。

数据集结构

数据集包含多个配置,每个配置对应不同的评估任务。以下是部分配置的详细信息:

  • 配置名称: harness_arc_challenge_25

    • 数据文件:
      • 分割: 2023_08_16T10_47_02.037059
        • 路径: **/details_harness|arc:challenge|25_2023-08-16T10:47:02.037059.parquet
      • 分割: latest
        • 路径: **/details_harness|arc:challenge|25_2023-08-16T10:47:02.037059.parquet
  • 配置名称: harness_hellaswag_10

    • 数据文件:
      • 分割: 2023_08_16T10_47_02.037059
        • 路径: **/details_harness|hellaswag|10_2023-08-16T10:47:02.037059.parquet
      • 分割: latest
        • 路径: **/details_harness|hellaswag|10_2023-08-16T10:47:02.037059.parquet
  • 配置名称: harness_hendrycksTest_5

    • 数据文件:
      • 分割: 2023_08_16T10_47_02.037059
        • 路径:
          • **/details_harness|hendrycksTest-abstract_algebra|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-anatomy|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-astronomy|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-business_ethics|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-clinical_knowledge|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-college_biology|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-college_chemistry|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-college_computer_science|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-college_mathematics|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-college_medicine|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-college_physics|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-computer_security|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-conceptual_physics|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-econometrics|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-electrical_engineering|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-elementary_mathematics|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-formal_logic|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-global_facts|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-high_school_biology|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-high_school_chemistry|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-high_school_computer_science|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-high_school_european_history|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-high_school_geography|5_2023-08-16T10:47:02.037059.parquet
          • **/details_harness|hendrycksTest-high_school_government_and_politics|5_2023-08-16T10:47:02.037059.parquet

最新结果

以下是 2023-08-16T10:47:02.037059 运行的最新结果

python { "all": { "acc": 0.4739285188775824, "acc_stderr": 0.035185125877572575, "acc_norm": 0.4774082437104984, "acc_norm_stderr": 0.035170487487277746, "mc1": 0.3329253365973072, "mc1_stderr": 0.016497402382012055, "mc2": 0.49276058409873585, "mc2_stderr": 0.01516224977207343 }, "harness|arc:challenge|25": { "acc": 0.523037542662116, "acc_stderr": 0.014595873205358269, "acc_norm": 0.537542662116041, "acc_norm_stderr": 0.014570144495075581 }, "harness|hellaswag|10": { "acc": 0.6015733917546305, "acc_stderr": 0.004885735963346904, "acc_norm": 0.7923720374427405, "acc_norm_stderr": 0.0040477996462346365 }, "harness|hendrycksTest-abstract_algebra|5": { "acc": 0.33, "acc_stderr": 0.04725815626252604, "acc_norm": 0.33, "acc_norm_stderr": 0.04725815626252604 }, "harness|hendrycksTest-anatomy|5": { "acc": 0.4888888888888889, "acc_stderr": 0.04318275491977976, "acc_norm": 0.4888888888888889, "acc_norm_stderr": 0.04318275491977976 }, "harness|hendrycksTest-astronomy|5": { "acc": 0.42105263157894735, "acc_stderr": 0.040179012759817494, "acc_norm": 0.42105263157894735, "acc_norm_stderr": 0.040179012759817494 }, "harness|hendrycksTest-business_ethics|5": { "acc": 0.48, "acc_stderr": 0.050211673156867795, "acc_norm": 0.48, "acc_norm_stderr": 0.050211673156867795 }, "harness|hendrycksTest-clinical_knowledge|5": { "acc": 0.5056603773584906, "acc_stderr": 0.030770900763851316, "acc_norm": 0.5056603773584906, "acc_norm_stderr": 0.030770900763851316 }, "harness|hendrycksTest-college_biology|5": { "acc": 0.5, "acc_stderr": 0.04181210050035455, "acc_norm": 0.5, "acc_norm_stderr": 0.04181210050035455 }, "harness|hendrycksTest-college_chemistry|5": { "acc": 0.3, "acc_stderr": 0.046056618647183814, "acc_norm": 0.3, "acc_norm_stderr": 0.046056618647183814 }, "harness|hendrycksTest-college_computer_science|5": { "acc": 0.39, "acc_stderr": 0.04902071300001975, "acc_norm": 0.39, "acc_norm_stderr": 0.04902071300001975 }, "harness|hendrycksTest-college_mathematics|5": { "acc": 0.31, "acc_stderr": 0.04648231987117316, "acc_norm": 0.31, "acc_norm_stderr": 0.04648231987117316 }, "harness|hendrycksTest-college_medicine|5": { "acc": 0.4161849710982659, "acc_stderr": 0.03758517775404947, "acc_norm": 0.4161849710982659, "acc_norm_stderr": 0.03758517775404947 }, "harness|hendrycksTest-college_physics|5": { "acc": 0.19607843137254902, "acc_stderr": 0.03950581861179962, "acc_norm": 0.19607843137254902, "acc_norm_stderr": 0.03950581861179962 }, "harness|hendrycksTest-computer_security|5": { "acc": 0.57, "acc_stderr": 0.049756985195624284, "acc_norm": 0.57, "acc_norm_stderr": 0.049756985195624284 }, "harness|hendrycksTest-conceptual_physics|5": { "acc": 0.4, "acc_stderr": 0.03202563076101735, "acc_norm": 0.4, "acc_norm_stderr": 0.03202563076101735 }, "harness|hendrycksTest-econometrics|5": { "acc": 0.2631578947368421, "acc_stderr": 0.041424397194893624, "acc_norm": 0.2631578947368421, "acc_norm_stderr": 0.041424397194893624 }, "harness|hendrycksTest-electrical_engineering|5": { "acc": 0.43448275862068964, "acc_stderr":

搜集汇总
数据集介绍
open-llm-leaderboard/details_edor__Hermes-Platypus2-mini-7B 数据集图片
构建方式
在大型语言模型评估领域,对模型性能进行系统化、标准化的评测至关重要。该数据集是Hugging Face Open LLM Leaderboard在评测edor/Hermes-Platypus2-mini-7B模型过程中自动生成的产物。其构建方式基于一次完整的评测运行(2023年8月16日),涵盖了61个不同的评测任务配置,每个配置对应一个独立的评估任务,如ARC挑战集、HellaSwag以及涵盖数十个学科领域的MMLU(HendrycksTest)等。数据以Parquet格式存储,每个任务配置下包含以时间戳命名的运行分割和指向最新结果的'train'分割,同时设有一个专门的'results'配置来汇总所有任务的聚合指标,为模型性能的横向对比提供了结构化的数据基础。
使用方法
研究者可通过Hugging Face的datasets库便捷地加载该数据集。例如,使用load_dataset函数,指定数据集名称'open-llm-leaderboard/details_edor__Hermes-Platypus2-mini-7B'及目标任务配置(如'harness_truthfulqa_mc_0'),并选择'train'分割即可获取最新结果。若要回溯历史版本,则可通过具体时间戳命名的分割(如'2023_08_16T10_47_02.037059')加载对应运行的数据。加载后的数据可作为DataFrame进行查询与分析,用于复现排行榜指标、比较模型在不同学科上的能力差异,或作为后续模型微调与评估的基准参考,充分满足科研与工程场景下的多样化需求。
背景与挑战
背景概述
在大规模语言模型(LLM)性能评估领域,Open LLM Leaderboard由Hugging Face于2023年发起,旨在为开源社区提供一个标准化、可复现的模型评测平台。该数据集是围绕模型edor/Hermes-Platypus2-mini-7B在排行榜上的单次评估运行而自动构建的,创建时间为2023年8月16日,主要研究人员来自Hugging Face团队。其核心研究问题在于如何系统性地量化一个7B参数级别模型在多维度任务上的表现,涵盖常识推理、学术知识、伦理判断等领域。该数据集通过记录61个配置下的详细结果,为社区提供了透明、细粒度的模型能力快照,推动了开源LLM的公平比较与迭代优化,对后续模型开发与基准测试研究产生了重要影响。
当前挑战
该数据集所解决的领域问题是大规模语言模型的标准化性能评估,其挑战主要体现在两个方面。其一,模型评估本身面临任务多样性与指标一致性难题,例如ARC Challenge、HellaSwag、MMLU等任务分别测试不同认知能力,如何在同一框架下公平聚合并解释acc、mc1、mc2等指标,避免因任务难度差异导致的偏差,是核心挑战。其二,构建过程中需应对数据自动采集与版本管理的复杂性,每次评估运行生成独立的时间戳分片,确保结果可追溯与可复现,同时维护“latest”分片以反映最新状态,这对数据管道的一致性和存储架构提出了较高要求。
常用场景
经典使用场景
在大型语言模型(LLM)的蓬勃发展中,如何系统性地评估模型的多维能力成为关键议题。该数据集专为Open LLM Leaderboard的评估流程而生,将模型edor/Hermes-Platypus2-mini-7B在61个任务上的推理结果以结构化形式存储,涵盖ARC挑战赛、HellaSwag常识推理以及涵盖57个学科的MMLU测试等经典基准。研究者可通过加载特定配置与分割,复现模型在各项任务中的精确表现,从而进行横向对比与纵向追踪。这一设计不仅为模型开发者提供了便捷的性能诊断工具,更为社区构建了透明、可复现的评估标准,推动了LLM评测范式的规范化进程。
解决学术问题
该数据集直面大模型评估中结果碎片化与不可复现的痛点。传统上,模型评测结果散见于论文或博客,缺乏统一格式,难以进行系统比较与元分析。通过将每次运行的详细得分(包括准确率、标准化准确率、标准误等)整合为标准化数据集,它解决了跨模型、跨任务性能对比的学术难题。例如,研究者可借此量化Hermes-Platypus2-mini-7B在TruthfulQA上的诚实性(mc1为33.29%)与在HellaSwag上的常识推理能力(acc_norm达79.24%),从而揭示模型在不同认知维度上的优势与短板。这一数据基础设施的建立,为理解模型泛化边界、诊断失败案例提供了坚实的数据支撑,显著提升了LLM研究的科学性与严谨性。
实际应用
在实际应用中,该数据集扮演着模型选型与部署决策的“裁判员”角色。企业或研究机构在挑选适合特定场景的LLM时,可依据数据集中的细粒度结果进行权衡。例如,若需构建教育领域的问答系统,可优先关注MMLU子任务中高中生物(52.26%)、心理学(64.40%)等科目的得分;若面向事实性问答,则需审视TruthfulQA的mc2指标(49.28%)。此外,数据集的时间戳分割支持追踪模型迭代过程中的性能演变,为模型微调策略的优化提供反馈闭环。这种数据驱动的评估方式,加速了从实验室模型到工业级应用的转化进程。
数据集最近研究
最新研究方向
在大型语言模型(LLM)评估领域,Open LLM Leaderboard 已成为衡量模型综合性能的权威基准。Hermes-Platypus2-mini-7B 作为参数量为7B的轻量级模型,其评估数据集聚焦于多任务、多领域的知识推理与事实一致性检测。该数据集涵盖 ARC-Challenge、HellaSwag 等常识推理任务,以及涵盖57个学科的 MMLU 基准,并引入 TruthfulQA 以考察模型生成内容的真实性。前沿研究方向集中于通过细粒度评估揭示小参数模型在跨领域泛化与事实性上的能力边界,例如在抽象代数、大学物理等专业任务上表现波动,而在高中世界历史、市场营销等任务上展现较高准确率。这一评估体系不仅为模型迭代提供量化反馈,更推动了开源社区对高效、轻量级 LLM 实用性的深入探索,其意义在于为资源受限场景下的模型部署提供可靠选型依据。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务