WinoQueer-NL
收藏资源简介:
WinoQueer-NL是面向荷兰语语言模型的LGBTQ+偏见评估数据集,由马斯特里赫特大学基于英文WinoQueer基准创建,经文化适配与社区验证而成。该数据集包含42,906对刻板与反刻板句子,涵盖167个预测谓词、60个荷兰语名字及9种性别身份,总计约34万token。通过翻译、机器辅助校正及43名荷兰酷儿参与者的在线调查构建,确认145个文化相关刻板印象并新增22个荷兰特有偏见。其旨在评估荷兰语及多语言模型对LGBTQ+群体的偏见,揭示模型在跨性别身份上偏见高达97%的显著差异,为公平性研究提供关键基准。
WinoQueer-NL is an LGBTQ+ bias evaluation dataset for Dutch language models, developed by Maastricht University based on the English WinoQueer benchmark, with cultural adaptation and community validation. This dataset contains 42,906 pairs of stereotypical and anti-stereotypical sentences, covering 167 predictive predicates, 60 Dutch names, and 9 gender identities, totaling approximately 340,000 tokens. It was constructed via translation, machine-assisted correction, and an online survey with 43 Dutch queer participants. During this process, 145 culturally relevant stereotypes were confirmed, and 22 Dutch-specific biases were newly added. This dataset aims to evaluate bias against LGBTQ+ groups in Dutch and multilingual models, revealing a significant disparity with up to 97% bias against transgender identities, providing a critical benchmark for fairness research.




