RhaegarTargaryen/BenMRXETB11002
收藏资源简介:
Open-MM-RL是一个多模态STEM推理数据集,涵盖物理学、数学、生物学和化学领域。它专为需要模型解释视觉信息并结合逐步分析推理的问题而设计。与现有的多模态推理基准相比,Open-MM-RL通过包含多面板和多图像任务来扩展评估设置,这些任务需要整合更复杂的视觉上下文,因为现实生活中的问题很少局限于单一图像,而是信息通常分散在多个相关图像中,要求科学工作者跨图像推理以找到解决方案。数据集包括三种多模态输入格式:单图像问题(一张图像配一个问题)、多面板问题(复合或基于面板的视觉配一个问题)和多图像问题(多张独立图像配一个问题)。这些格式通过要求模型不仅从文本推理,还要跨视觉布局、多个视图和分布式证据推理,增加了任务复杂性。所有格式的问题都构建为自包含、无歧义、推理密集和可验证的,使得数据集既可作为评估基准,也可作为专注于推理的模型的训练资源。该数据集的一个关键区别特征是它专注于所有三种多模态格式的博士级别STEM问题解决,这使得评估高级主题推理和模型跨日益复杂视觉输入综合信息的能力成为可能。与依赖标题的科学图形基准不同,该数据集中的示例设计为直接从提供的图像和问题中回答。
Open-MM-RL is a multimodal STEM reasoning dataset covering Physics, Mathematics, Biology, and Chemistry. It is designed for problems that require models to interpret visual information and combine it with step-by-step analytical reasoning. Compared with existing multimodal reasoning benchmarks, Open-MM-RL broadens the evaluation setting beyond standard single-image question answering by including multi-panel and multi-image tasks that require integrating information across more complex visual contexts. As rarely in real-life problems is context confined to a single image. Instead, the necessary information is often fragmented across multiple related images, requiring scientists to reason across them to find the solution. The dataset includes three multimodal input formats: Single-image problems (one image paired with one question), Multi-panel problems (a composite or panel-based visual paired with one question), and Multi-image problems (multiple separate images paired with one question). These formats increase task complexity by requiring models to reason not only from text, but also across visual layouts, multiple views, and distributed evidence. Across all formats, problems are constructed to be self-contained, unambiguous, reasoning-intensive, and verifiable making the dataset useful both as an evaluation benchmark and as a training resource for reasoning-focused models. A key distinguishing feature of this dataset is its focus on PhD-level STEM problem solving across all three multimodal formats. This makes it possible to assess both advanced subject-matter reasoning and a models ability to synthesize information across increasingly complex visual inputs. Unlike scientific figure benchmarks that rely significantly on captions, examples in this dataset are designed to be answered directly from the provided image or images together with the question.




