SYMBALBENCH
收藏资源简介:
SYMBALBENCH是由斯坦福大学研究团队构建的基准数据集,旨在系统评估多模态大语言模型生成图像描述中的系统性错位检测方法。该数据集规模庞大,包含170万张图像-文本对,涵盖自然图像和医学影像两大领域,并进一步细分为420个视觉-语言子集,每个子集均标注了系统性错位真值。数据来源于多模态大语言模型生成的描述文本,通过结构化构建过程确保标注质量。该数据集主要应用于多模态人工智能的可靠性评估领域,旨在揭示和诊断模型在生成描述时存在的系统性偏见与错误关联,为提升模型安全性与可解释性提供关键基准。
SYMBALBENCH is a benchmark dataset developed by a research team from Stanford University, aiming to systematically evaluate methods for detecting systematic misalignment in image captions generated by multimodal large language models. Boasting a large scale, the dataset comprises 1.7 million image-text pairs spanning two domains: natural images and medical imaging, and is further subdivided into 420 vision-language subsets, each annotated with ground truth labels for systematic misalignment. The data is sourced from caption texts generated by multimodal large language models, with its annotation quality guaranteed via a structured construction workflow. This dataset is primarily utilized in the field of reliability assessment for multimodal AI, with the goal of uncovering and diagnosing systematic biases and erroneous associations in model-generated captions, thereby providing a pivotal benchmark for enhancing model safety and interpretability.

- 1Symbal: Detecting Systematic Misalignments in Model-Generated Captions斯坦福大学; HOPPR · 2026年




