遇见数据集

agcbench-2026/AGC-Judge-Training-Data

收藏
Hugging Face2026-05-08 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是AGC-Judge(开放权重评分器)的训练数据,用于AGC-Bench(人工通用创造力基准)的训练和验证。数据集包含聊天格式的监督和评估分割,每条记录是一个三消息对话:1. system:AGC-Judge的评分指令;2. user:基准评分标准、基准提示和待评分的模型响应;3. assistant:作为黄金目标的JRT校正整数分数。数据集提供多个配置:messages(对话数据)、metadata(元数据,包括基准、模型、项目ID、指标、分数等)和manifest(清单文件)。分割包括训练集(48,299行,过滤后的监督数据)、验证集(6,883行,过滤后的验证监督数据)、测试集(14,457行,分布内保留示例)、holdout_models(10,966行,训练中未见过的模型)和holdout_benches(8,803行,训练中未见过的基准/评分标准)。数据集适用于聊天微调管道,支持直接加载或使用压缩JSONL文件。数据集基于多个源基准,用户需遵守相关许可证和条款。

This dataset contains the chat-format supervision and evaluation splits used to train and validate AGC-Judge, the open-weight scorer released with AGC-Bench (Artificial General Creativity Benchmark). Each row in the messages config is a three-message chat conversation: 1. system: scoring instruction for AGC-Judge; 2. user: benchmark rubric, benchmark prompt, and model response to score; 3. assistant: the JRT-corrected integer score used as the gold target. The dataset includes multiple configs: messages (conversation data), metadata (provenance data such as benchmark, model, item_id, metric, score, etc.), and manifest (manifest file). Splits consist of train (48,299 rows, filtered supervision data), validation (6,883 rows, filtered validation supervision), test (14,457 rows, in-distribution held-out examples), holdout_models (10,966 rows, held-out models unseen during fine-tuning), and holdout_benches (8,803 rows, held-out benchmarks/rubrics unseen during fine-tuning). The dataset is usable by chat fine-tuning pipelines that accept OpenAI-style messages records and is released for research transparency, with users responsible for respecting source-benchmark licenses and terms.

提供机构:
agcbench-2026
二维码
社区交流群
二维码
科研交流群
商业服务