SciVideoBench
收藏资源简介:
SciVideoBench是一个专门设计用于评估大型多模态模型在科学领域的高级视频推理能力的严格基准。它由1000个精心设计的多项选择题组成,这些问题来自涵盖超过25个专业学术领域的尖端科学实验视频。每个问题都需要复杂的领域特定知识、精确的时空感知和复杂的逻辑推理,有效地挑战模型的高级认知能力。SciVideoBench的数据集来源于超过25个不同的学术领域,包括流体力学、分析化学、神经科学和肿瘤学等。数据集的内容包括来自物理学、化学、生物学和医学等领域的241个研究级实验视频。这些视频来自Journal of Visualized Experiments (JoVE)平台,是一个同行评审的平台,专注于发表跨广泛科学学科的方法论视频。SciVideoBench的数据集创建过程采用多阶段、人机协作的流程,包括挖掘相关实验手稿、利用大型多模态模型进行初步问题生成,以及让领域专家验证问题-答案对,并筛选掉无法回答或与视频无关的问题。SciVideoBench旨在解决科学领域中的复杂视频推理问题,推动大型多模态模型在视频推理能力方面的进步。
SciVideoBench is a rigorous benchmark specifically designed to evaluate the advanced video reasoning capabilities of large multimodal models in the scientific domain. It consists of 1,000 carefully crafted multiple-choice questions sourced from cutting-edge scientific experiment videos spanning over 25 specialized academic fields. Each question requires complex domain-specific knowledge, precise spatio-temporal perception, and sophisticated logical reasoning, effectively challenging the high-level cognitive abilities of the models. The dataset of SciVideoBench is sourced from more than 25 distinct academic disciplines, including fluid mechanics, analytical chemistry, neuroscience, oncology, and others. The dataset content includes 241 research-grade experimental videos from fields such as physics, chemistry, biology, and medicine. These videos are sourced from the Journal of Visualized Experiments (JoVE) platform, a peer-reviewed platform dedicated to publishing methodological videos across a wide range of scientific disciplines. The dataset creation process of SciVideoBench adopts a multi-stage, human-machine collaborative workflow, which includes mining relevant experimental manuscripts, generating initial questions using large multimodal models, and having domain experts validate the question-answer pairs while filtering out unanswerable or video-irrelevant questions. SciVideoBench aims to address complex video reasoning problems in the scientific field and promote the advancement of video reasoning capabilities for large multimodal models.

- 1通过中央佛罗里达大学, 北卡罗来纳大学教堂山分校, 斯坦福大学 · 2025年



