EuroGEST
收藏资源简介:
EuroGEST是一个包含71000个句子的数据集,这些句子与16个性别刻板印象相关联,覆盖了30种欧洲语言。该数据集扩展了现有的专家信息基准,涵盖了16个性别刻板印象,并使用翻译工具、质量评估指标和形态学启发式方法进行了扩展。人工评估证实,我们的数据生成方法在翻译和性别标签的准确性方面取得了高精度。我们使用EuroGEST评估了来自六个模型家族的24个多语言语言模型,结果表明,所有模型中所有语言中最强烈的刻板印象是女性美丽、富有同情心、整洁,而男性则是领导者、坚强、坚韧和专业。我们还表明,更大的模型编码性别刻板印象更强烈,而指令微调并不能始终如一地减少性别刻板印象。
EuroGEST is a dataset containing 71,000 sentences associated with 16 gender stereotypes, covering 30 European languages. It expands upon existing expert-curated benchmarks for 16 gender stereotypes, and was constructed using translation tools, quality assessment metrics, and morphological heuristics. Human evaluations confirm that our data generation method achieves high accuracy in both translation quality and gender labeling. We evaluated 24 multilingual large language models from six model families using EuroGEST, and the results demonstrate that across all languages and all models, the most salient gender stereotypes are that women are perceived as beautiful, compassionate, and tidy, while men are framed as leaders, strong, resilient, and professional. We further show that larger models encode gender stereotypes more strongly, and that instruction tuning does not consistently reduce gender stereotypes.




