rebase_vgs_gpt-oss-20b_lcb_v6_ns8_md4_bt0_1_seed42_lcb_v6_vgs_2gpu
收藏资源简介:
该数据集包含1,048个测试样本,每个样本包含问题、生成ID、生成内容、令牌数量、奖励值、问题索引、目标、任务、vf预测值和级别等多个特征。数据集主要用于评估生成模型的性能,包含丰富的评估指标如pass@1、pass@2等通过率指标,以及各种令牌统计信息。数据集中还记录了不同级别(1-4级)的策略输出令牌数和时间指标,可用于深入分析模型在不同难度级别下的表现。
This dataset contains 1,048 test samples, each of which includes multiple features such as question, generation ID, generated content, token count, reward value, question index, target, task, vf prediction value and level. It is mainly used to evaluate the performance of generative models, and includes rich evaluation metrics such as pass@1, pass@2 and other pass-rate indicators, as well as various token statistics. The dataset also records the number of policy output tokens and time metrics for different levels (levels 1-4), which can be used to conduct in-depth analysis of model performance across varying difficulty levels.
根据您提供的数据集详情页面内容,以下是对该数据集的总结:
数据集概述
该数据集名为 rebase_vgs_gpt-oss-20b_lcb_v6_ns8_md4_bt0_1_seed42_lcb_v6_vgs_2gpu,由用户 anirudhb11 上传至 Hugging Face。
数据特征
数据集包含以下 10 个字段:
- question:字符串类型,存储问题文本。
- generation_id:整数类型,生成结果的唯一标识。
- generation:字符串类型,模型生成的内容。
- num_tokens:整数类型,生成的 token 数量。
- reward:浮点数类型,奖励值。
- question_index:整数类型,问题索引。
- target:字符串类型,目标答案或参考。
- task:字符串类型,任务名称。
- vf_prediction:浮点数类型,价值函数预测值。
- level:整数类型,级别或难度等级。
数据划分
- 唯一数据划分:
test集 - 测试集规模:共包含 1048 个样本
- 数据集大小:压缩后约 916.95 MB,解压后约 1.57 GB
聚合指标
该数据集来源于某次生成实验,聚合了来自 16 个分片(shard) 的指标。关键性能指标如下:
- 生成质量指标:
pass@1:0.6355pass@2:0.6897pass@4:0.7156pass@8:0.7252maj@1至maj@8:均为 0
- Token 统计:
avg_response_tokens:5915.43median_response_tokens:4201.36total_generated_output_tokens:385,598
- 时间消耗:
generation_phase_time_s:390.93 秒total_time_s:666.75 秒
- 加权最佳指标:
w_best@1:0.6305w_best@2:0.6611w_best@4:0.6534w_best@8:0.6641
配置信息
- 配置名称:
default - 数据文件路径:
data/test-*(匹配多个分片文件)




