Bias in the Picture Benchmark
收藏资源简介:
包含1,343个新闻来源的图像-问题对,标注了人口统计和社会属性,用于评估视觉语言模型对包含年龄、性别、种族和职业等社会线索的真实新闻图像的响应
This dataset comprises 1,343 image-question pairs derived from news outlets, annotated with demographic and social attributes. It is designed to evaluate the responses of vision-language models to real news images containing social cues such as age, gender, race, occupation, and other similar social cues.
Bias in the Picture Benchmark 数据集概述
数据集简介
Bias in the Picture Benchmark 是一个用于评估视觉语言模型在包含社会线索的真实新闻图像中表现偏差的基准数据集。该数据集包含1,343个源自新闻的图像-问题对,标注了人口统计和社会属性。
核心特征
- 真实世界社会线索基准:包含1,343个新闻衍生的图像-问题对
- 多模态评估流程:提供运行VLM推理、清理输出和计算评估指标的工具
- 广泛模型覆盖:支持Aya Vision、Gemini/Gemma、Phi、Qwen2.5-VL、LLaMA、Molmo、CogVLM2、Paligemma、LLaVA、JanusPro等模型
- LLM-as-Judge评分:使用结构化评分标准系统评估准确性、偏差和忠实度
评估指标
- LLM-as-Judge指标:偏差、忠实度、相关性
- 标准NLP指标:BERTScore、METEOR、FrugalScore
数据集结构
- 数据文件:data/data.parquet
- 源代码:src/目录包含完整流程(数据准备、推理、指标计算)
- 文档:docs/目录包含使用指南
使用方式
- 安装依赖环境
- 运行处理流程:
- 数据预处理
- 模型推理
- 计算评估指标
引用信息
如需使用该基准数据集,请引用相关论文:
@misc{narayanan2025biaspicturebenchmarkingvlms, title={Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment}, author={Aravind Narayanan and Vahid Reza Khazaie and Shaina Raza}, year={2025}, eprint={2509.19659}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2509.19659} }
联系方式
- 问题反馈:https://github.com/VectorInstitute/bias-in-the-picture-benchmark/issues
- 作者邮箱:aravind.narayanan@vectorinstitute.ai




