ENSEONG/full-math-private-n256-Llama-3.2-3B-Instruct-bon
收藏资源简介:
该数据集是一个数学问题评估数据集,包含多个配置版本,每个配置基于不同的生成参数(如温度T=0.1、0.2、0.3,top_p=1.0,生成数量n=256,随机种子seed=0、42、64、128、256、512,聚合策略agg_strategy=last)生成。每个示例包含数学问题(problem)、难度等级(level)、类型(type)、标准解决方案(solution)、答案(answer)、多个生成完成(completions)、预测结果(pred、preds)、完成令牌数(completion_tokens)以及从1到256不同规模下的多数预测(pred_maj@*)和通过率(pass@*)。数据集旨在评估语言模型在数学问题上的性能,通过多参数设置生成多样化的预测和评估指标,适用于数学推理、模型评估和基准测试任务。
This dataset is a mathematical problem evaluation dataset containing multiple configuration versions, each generated based on different parameters (e.g., temperature T=0.1, 0.2, 0.3, top_p=1.0, number of generations n=256, random seeds seed=0, 42, 64, 128, 256, 512, aggregation strategy agg_strategy=last). Each example includes a mathematical problem, difficulty level, problem type, standard solution, answer, multiple generated completions, prediction results (pred, preds), completion token counts, and majority predictions (pred_maj@*) and pass rates (pass@*) for scales from 1 to 256. The dataset is designed to evaluate the performance of language models on mathematical problems, generating diverse predictions and evaluation metrics through multi-parameter settings, suitable for mathematical reasoning, model evaluation, and benchmarking tasks.




