weighted_rd_results
收藏资源简介:
该数据集是在模型`erinya/weighted_rd_results`的评估运行期间自动创建的,用于存储和追踪该模型在多个标准基准测试任务上的评估结果。数据集结构围绕评估运行组织,包含多个配置,每个配置对应一个被评估的任务,并且每个配置下有多个以运行时间戳命名的数据分割,其中train分割始终指向最新的评估结果。此外,一个名为results的专门配置汇总了所有运行的聚合评估数据。数据集内容包括模型在ARC(Easy和Challenge版本)、MMLU(涵盖抽象代数、解剖学、天文学、生物学、化学、计算机科学、数学、物理学、统计学、经济学、历史学、法学、哲学、医学、商学等数十个子领域)、GSM8K(数学问题求解)和WikiText(语言建模)等一系列广泛任务上的性能指标。这些指标通常包括每个任务的测试样本数量、原始准确率、标准化准确率以及相应的标准误差,对于GSM8K和WikiText还有特定的匹配度或困惑度指标。数据集的主要用途是记录、比较和分析语言模型在多样化评估任务上的性能表现,适用于模型评估、基准测试研究和性能趋势追踪等场景。
This dataset was automatically created during the evaluation run of the model `erinya/weighted_rd_results`. It is used to store and track the evaluation results of this model on multiple standard benchmark tasks. The dataset structure is organized around evaluation runs: it contains multiple configurations (each corresponding to an evaluated task), and each configuration includes multiple data splits named after run timestamps, where the train split always points to the latest evaluation results. Additionally, a specialized configuration named results aggregates evaluation data from all runs. The dataset content specifically includes performance metrics of the model on a wide range of tasks such as ARC (Easy and Challenge versions), MMLU (covering dozens of subfields including abstract algebra, anatomy, astronomy, biology, chemistry, computer science, mathematics, physics, statistics, economics, history, law, philosophy, medicine, business, etc.), GSM8K (mathematical problem-solving), and WikiText (language modeling). These metrics typically include the number of test samples per task (sample_len), raw accuracy (acc,none), normalized accuracy (acc_norm,none), and corresponding standard errors (stderr), with specific matching or perplexity metrics for GSM8K and WikiText. The primary purpose of this dataset is to record, compare, and analyze the performance of language models on diverse evaluation tasks, suitable for scenarios such as model evaluation, benchmark testing research, and performance trend tracking.
数据集概览
- 数据集名称:
erinya/weighted_rd_results - 数据集描述: 该数据集是在对模型
erinya/weighted_rd_results进行评测运行时自动创建的,用于存储各任务的评测结果。
数据集结构
- 配置: 数据集包含 0 个任务配置,每个配置对应一个被评测的任务。
- 运行记录: 数据集由 3 次运行产生,每次运行在各自的配置中以时间戳命名的 split 形式存在,
trainsplit 始终指向最新的结果。 - 聚合结果: 额外的配置
results存储所有运行的整体聚合结果。
最新运行结果
最新运行的时间戳为 2026-05-03T16-48-11.185306,结果文件位于 erinya/weighted_rd_results/results_2026-05-03T16-48-11.185306.json。以下是该运行在各项任务上的核心指标:
ARCeasy
- 样本数: 2376
- 准确率(acc): 0.3737
- 标准化准确率(acc_norm): 0.3523
ARCChallenge
- 样本数: 1172
- 准确率(acc): 0.2961
- 标准化准确率(acc_norm): 0.3131
MMLU 各子任务 (部分)
MMLU 子任务均使用 acc,none 作为准确率评估指标:
- abstract_algebra: 准确率 0.1900 (样本数: 100)
- anatomy: 准确率 0.1852 (样本数: 135)
- astronomy: 准确率 0.2368 (样本数: 152)
- college_biology: 准确率 0.2778 (样本数: 144)
- college_chemistry: 准确率 0.1900 (样本数: 100)
- college_computer_science: 准确率 0.2500 (样本数: 100)
- college_mathematics: 准确率 0.2100 (样本数: 100)
- college_physics: 准确率 0.2157 (样本数: 102)
- computer_security: 准确率 0.2800 (样本数: 100)
- conceptual_physics: 准确率 0.2638 (样本数: 235)
- electrical_engineering: 准确率 0.2690 (样本数: 145)
- elementary_mathematics: 准确率 0.2116 (样本数: 378)
- high_school_biology: 准确率 0.1968 (样本数: 310)
- high_school_chemistry: 准确率 0.1626 (样本数: 203)
- high_school_computer_science: 准确率 0.2500 (样本数: 100)
- high_school_mathematics: 准确率 0.2111 (样本数: 270)
- high_school_physics: 准确率 0.2053 (样本数: 151)
- high_school_statistics: 准确率 0.1528 (样本数: 216)
- machine_learning: 准确率 0.2946 (样本数: 112)
- business_ethics: 准确率 0.3000 (样本数: 100)
- clinical_knowledge: 准确率 0.2113 (样本数: 265)
- college_medicine: 准确率 0.2023 (样本数: 173)
- global_facts: 准确率 0.1700 (样本数: 100)
- human_aging: 准确率 0.3139 (样本数: 223)
- management: 准确率 0.1748 (样本数: 103)
- marketing: 准确率 0.2906 (样本数: 234)
- medical_genetics: 准确率 0.3200 (样本数: 100)
- miscellaneous: 准确率 0.2363 (样本数: 783)
- nutrition: 准确率 0.2157 (样本数: 306)
- professional_accounting: 准确率 0.2376 (样本数: 282)
- professional_medicine: 准确率 0.1985 (样本数: 272)
- virology: 准确率 0.2831 (样本数: 166)
- econometrics: 准确率 0.2281 (样本数: 114)
- high_school_geography: 准确率 0.1768 (样本数: 198)
- high_school_government_and_politics: 准确率 0.2332 (样本数: 193)
- high_school_macroeconomics: 准确率 0.2308 (样本数: 390)
- high_school_microeconomics: 准确率 0.2353 (样本数: 238)
- high_school_psychology: 准确率 0.1908 (样本数: 545)
- human_sexuality: 准确率 0.2595 (样本数: 131)
- professional_psychology: 准确率 0.2500 (样本数: 612)
- public_relations: 准确率 0.2182 (样本数: 110)
- security_studies: 准确率 0.1878 (样本数: 245)
- sociology: 准确率 0.2488 (样本数: 201)
- us_foreign_policy: 准确率 0.2700 (样本数: 100)
- formal_logic: 准确率 0.2540 (样本数: 126)
- high_school_european_history: 准确率 0.2424 (样本数: 165)
- high_school_us_history: 准确率 0.3186 (样本数: 204)
- high_school_world_history: 准确率 0.3333 (样本数: 237)
- international_law: 准确率 0.2314 (样本数: 121)
- jurisprudence: 准确率 0.2593 (样本数: 108)
- logical_fallacies: 准确率 0.2209 (样本数: 163)
- moral_disputes: 准确率 0.2486 (样本数: 346)
- moral_scenarios: 准确率 0.2380 (样本数: 895)
- philosophy: 准确率 0.1865 (样本数: 311)
- prehistory: 准确率 0.2130 (样本数: 324)
- professional_law: 准确率 0.2503 (样本数: 1534)
- world_religions: 准确率 0.3216 (样本数: 171)
GSM8K
- 样本数: 1319
- 严格匹配(exact_match,strict-match): 0.0
- 灵活提取匹配(exact_match,flexible-extract): 0.1289
Wikitext
- 样本数: 62
- 词级困惑度(word_perplexity): 142.7229
- 字节级困惑度(byte_perplexity): 2.5287
- 比特每字节(bits_per_byte): 1.3384
数据加载示例
可通过以下代码加载最新结果: python from datasets import load_dataset data = load_dataset( "erinya/weighted_rd_results", name="erinya/weighted_rd_results__results", split="latest" )




