thoughtworks/CulturalCounterfactuals
收藏资源简介:
Cultural Counterfactuals是一个高质量的合成图像数据集,用于测量大型视觉语言模型(LVLM)中的文化偏见。它包含59,827张图像,分为10,331个反事实集,涵盖宗教、国籍和社会经济地位三个文化维度。在每个反事实集中,同一个合成个体被描绘在不同的文化背景下(例如,同一个人站在基督教教堂、清真寺或犹太教堂前),从而可以控制测量LVLM输出如何仅随文化背景变化。数据集还提供了详细的构建过程、文件布局、快速开始指南、许可证信息和引用方式。
Cultural Counterfactuals is a high-quality synthetic image dataset developed to measure cultural bias in large vision-language models (LVLMs). It contains 59,827 images divided into 10,331 counterfactual sets, covering three cultural dimensions: religion, nationality, and socioeconomic status. In each counterfactual set, the same synthetic individual is depicted under different cultural contexts (e.g., the same person standing in front of a Christian church, a mosque, or a synagogue), thus enabling controlled measurement of how LVLM outputs change solely with variations in cultural background. The dataset also provides detailed construction procedures, file layout, quick start guide, license information, and citation methods.




