MSCoRe
收藏资源简介:
MSCoRe是一个用于评估大型语言模型在复杂、多阶段场景中协同推理能力的新基准。该数据集包含来自汽车、医药、电子和能源领域的126,696个特定领域的问答实例。数据集的创建采用了动态采样、迭代问答生成和多级质量评估的流程,以确保数据质量。任务被细分为三个难度级别:简单、中等和困难,以便进行细粒度的分析。MSCoRe为社区提供了一个有价值的资源,用于评估和改进大型语言模型的多阶段推理能力。
MSCoRe is a novel benchmark for evaluating the collaborative reasoning capabilities of large language models (LLMs) in complex, multi-stage scenarios. This dataset includes 126,696 domain-specific question-answering instances across the automotive, pharmaceutical, electronics, and energy sectors. The dataset was developed through a workflow involving dynamic sampling, iterative question-answering generation, and multi-level quality assessment to ensure data quality. The tasks are divided into three difficulty levels: easy, medium, and hard, to enable fine-grained analysis. MSCoRe offers a valuable resource for the research community to evaluate and enhance the multi-stage reasoning capabilities of large language models.




