MathArena/arxivmath-training_outputs
收藏资源简介:
该数据集包含从过去的ArXiv文章中生成的训练数据,以及由Qwen3.6-35B模型生成的输出。由于一个错误,并非所有来自MathArena/arxivmath-training数据集的行都获得了输出。数据集包含多个字段,如问题索引、问题陈述、模型名称、模型配置、尝试索引、完整对话记录、用户提示、模型回答、输入和输出令牌数、成本估计、来源标识、黄金答案、解析答案和正确性判断等。这些数据用于评估大型语言模型在数学问题上的表现,并支持自动评分和成本分析。
This dataset contains training data generated from past arXiv articles, as well as outputs produced by the Qwen3.6-35B model. Due to an existing error, not all rows from the MathArena/arxivmath-training dataset have obtained corresponding outputs. The dataset includes multiple fields, such as question index, problem statement, model name, model configuration, attempt index, full conversation history, user prompt, model response, number of input and output tokens, cost estimate, source identifier, gold answer, parsed answer, and correctness judgment. This data is utilized to evaluate the performance of large language models on mathematical problems, and supports automatic scoring and cost analysis.




