preprocessed-full-math-private-Qwen2.5-3B-Instruct-bon
收藏资源简介:
该数据集包含多个配置版本的数学问题解答评估数据,每个配置包含5000个训练样本。核心字段包括数学问题描述(problem)、难度级别(level)、问题类型(type)、标准解答(solution)和答案(answer)。数据集特别关注模型预测性能评估,包含64种不同参数组合下的预测结果(completions)及对应的评分(scores),以及加权预测(pred_weighted)、多数表决预测(pred_maj)和朴素预测(pred_naive)等多种预测策略在不同样本量下的结果(@1到@64)。评估指标包含通过率(pass@k)和正确性判断(is_correct)。数据预处理元信息记录了预测字段数量、处理时间和版本号。该数据集适用于数学自动解题模型的性能评估和预测策略比较研究。
This dataset contains evaluation data for mathematical problem-solving assessments across multiple configuration versions, with 5,000 training samples per configuration. The core fields include the mathematical problem description (problem), difficulty level (level), problem type (type), standard solution (solution), and final answer (answer). Specifically focusing on model prediction performance evaluation, the dataset includes prediction results (completions) and their corresponding scores under 64 distinct parameter combinations, as well as results of multiple prediction strategies such as weighted prediction (pred_weighted), majority voting prediction (pred_maj), and naive prediction (pred_naive) across different sample sizes (@1 to @64). The evaluation metrics cover pass rate (pass@k) and correctness judgment (is_correct). The data preprocessing metadata records the number of prediction fields, processing time, and version number. This dataset is applicable to performance evaluation of automated mathematical problem-solving models and comparative research on prediction strategies.



