nandansarkar/model_that_was_trained_on_50_samples_no_length_norm-now_trained_on_og-1k-data_eval_c693
收藏数据链接:
官方服务:
资源简介:
该数据集包含在四个不同基准上进行的预计算模型输出,用于评估:AIME24、AIME25、GPQADiamond和JEEBench。它详细提供了每个基准的准确度结果、标准差以及执行的运行次数。对于AIME24和AIME25,各有20次运行;而对于GPQADiamond和JEEBench,各有10次运行。文件还提供了每次运行解答的问题数和总问题数。
This dataset includes precomputed model outputs for evaluation on four different benchmarks: AIME24, AIME25, GPQADiamond, and JEEBench. It provides detailed accuracy results, standard deviations, and the number of runs performed for each benchmark. AIME24 and AIME25 each have 20 runs, while GPQADiamond and JEEBench have 10 runs each. The file also includes the number of questions solved and the total number of questions for each run.
提供机构:
nandansarkar


