StylisticBias
收藏资源简介:
StylisticBias是由慕尼黑工业大学等机构构建的用于评估多模态大语言模型属性级社会偏见的基准数据集。该数据集包含约25,000张合成人脸图像,基于500个基础人脸通过单属性编辑生成,涵盖皮肤特征、发型、服饰风格等约50种视觉属性,数据来源于Imagen 4和Gemini 2.5 Flash Image生成的合成图像。创建过程采用受控生成方法,在保持身份不变的前提下系统性地修改单一视觉属性,并通过人工验证确保图像质量。该数据集旨在探究特定视觉线索如何影响模型的社会判断,为多模态模型的细粒度偏见评估提供标准化测试平台,可应用于算法公平性研究和社会感知计算领域。
StylisticBias is a benchmark dataset developed by institutions including the Technical University of Munich for evaluating attribute-level social biases in multimodal large language models. This dataset contains approximately 25,000 synthetic facial images, which are generated via single-attribute editing based on 500 base facial portraits, covering around 50 visual attributes such as skin features, hairstyles, clothing styles and more. The data is sourced from synthetic images generated by Imagen 4 and Gemini 2.5 Flash Image. Its creation adopts a controlled generation framework that systematically modifies a single visual attribute while maintaining the identity of the original face, and manual verification is performed to guarantee image quality. This dataset is designed to investigate how specific visual cues influence the social judgments of models, providing a standardized testbed for fine-grained bias evaluation of multimodal models, and can be applied to the fields of algorithmic fairness research and social perception computing.




