Examining Gender and Racial Bias in Vision-language Models within Colonial Contexts: Dataset
收藏资源简介:
Vision-language models~(VLMs) are increasingly used for image retrieval in cultural heritage collections. Yet these models are trained and evaluated on modern web-scraped data which might cause unobserved biases when applied to historical imagery. While prior work has documented gender and racial biases in VLMs using contemporary datasets, no evaluation datasets exist for historical archives. In this paper, we address this gap by introducing a new dataset for evaluating demographic bias in VLMs on photograph collections from colonial contexts. Our dataset comprises annotated photographs for demographic queries spanning gender and skin tone. Using this dataset, we examine bias in multiple VLMs through retrieval performance across queries, measuring disparities using established information retrieval metrics. Experimental results show that gender biases documented in contemporary datasets persist on historical imagery, with all tested models retrieving certain demographic groups substantially better than others. Notably, models specifically trained for debiasing fail to mitigate bias on historical images and can introduce new racial retrieval disparities. These findings underscore the need for domain-specific debiasing approaches and caution against deploying off-the-shelf VLMs in archival contexts without proper evaluation.



