zlab-princeton/WorldBench
收藏资源简介:
WorldBench是一个视觉多样化的多模态推理基准,用于评估现代多模态大语言模型(MLLMs)是否能够在广泛的视觉世界中进行推理。该基准围绕一个涵盖七个领域的广泛视觉分类法组织:生物、物体、场景、数字世界、学术、文档/图表/表格和代理。这个Hugging Face数据集包含整合的WorldBench评估分割:2,000个精心策划的多选题视觉推理示例,每个示例都配有一张图像和元数据,描述了细粒度视觉类别和高级领域。数据集格式为parquet,包含图像、问题、答案选项、答案键、类别和领域字段。
WorldBench is a visually diverse multimodal reasoning benchmark for evaluating whether modern Multimodal Large Language Models (MLLMs) can reason across the breadth of the visual world. The benchmark is organized around a broad visual taxonomy spanning seven domains: Living Things, Objects, Scenes, Digital World, Academics, Documents/Charts/Tables, and Agents. This Hugging Face dataset contains the consolidated WorldBench evaluation split: 2,000 curated multiple-choice visual reasoning examples, each paired with an image and metadata describing the fine-grained category and high-level domain. The dataset is in a simple parquet format with fields for image, question, answer choices, answer key, category, and domain.




