ProJudgeBench 和 ProJudge-173k
收藏资源简介:
ProJudgeBench是一个专门为评估多模态大型语言模型(MLLM)作为过程评判员的能力而设计的全面基准。该数据集包含2400个测试案例和50118个步骤级别的标签,涵盖数学、物理、化学和生物学四个学科,难度多样,内容多模态。每个步骤都由人类专家精心标注正确性、错误类型和解释,使得可以对评判员检测、分类和诊断错误的能力进行系统评估。ProJudge-173k是一个大规模的指令微调数据集,旨在对步骤-by-step推理进行细致评估。
ProJudgeBench is a comprehensive benchmark specifically designed to evaluate the capability of multimodal large language models (MLLMs) as process judges. This dataset contains 2400 test cases and 50118 step-level labels, covering four disciplines including mathematics, physics, chemistry and biology, with diverse difficulty levels and multimodal content. Each step is meticulously annotated by human experts with correctness, error types and corresponding explanations, enabling systematic assessment of the judge's ability to detect, classify and diagnose errors. ProJudge-173k is a large-scale instruction tuning dataset aimed at detailed evaluation of step-by-step reasoning.




