laion/delphi-1e23-25b-stageE-rl-eval-artifacts
收藏资源简介:
该数据集包含delphi-1e23 25B Stage-E RL扫描的原始evalchemy评估输出,用于模型性能评估。具体包括5个模型(一个SFT基线模型wc50m和四个RL单元D1-D4)在两个任务(MATH500和gsm8k)上的评估结果。文件分为结果JSON和样本JSONL:结果文件记录聚合指标和完整运行配置(如MATH500的准确率和gsm8k的灵活/严格精确匹配),样本文件提供每个示例的详细信息,如问题、标准答案、模型输出、提取答案和正确性。任务设置中,MATH500采用0-shot,最大生成标记3072;gsm8k采用5-shot,最大生成标记2048。评估通过vLLM TP=6进行,使用delphi_v0聊天模板,并指定了相关参数。模型来源于laion组织。
This dataset contains raw evalchemy evaluation outputs for the delphi-1e23 25B Stage-E RL sweep, used for model performance assessment. It includes evaluation results for 5 models (an SFT baseline wc50m and four RL cells D1–D4) on two tasks (MATH500 and gsm8k). Files consist of result JSON and sample JSONL: result files record aggregate metrics and full run configurations (e.g., accuracy for MATH500 and exact_match flexible/strict for gsm8k), while sample files provide per-example details such as problem, gold answer, model output, extracted answer, and correctness. Task settings involve MATH500 with 0-shot and a maximum generation token limit of 3072, and gsm8k with 5-shot and a limit of 2048 tokens. Evaluation is conducted via vLLM TP=6 using the delphi_v0 chat template with specified parameters. Models are sourced from the laion organization.




