SeeGULL Multilingual
收藏资源简介:
SeeGULL Multilingual是由谷歌研究团队创建的一个大规模多语言数据集,包含超过25,000个社会刻板印象,涵盖20种语言和23个地区。该数据集通过结合大型语言模型生成和基于文化的验证方法构建,旨在解决多语言模型评估中的安全性和公平性问题。数据集内容包括身份群体的刻板印象及其在不同地区的冒犯性评级,用于帮助识别模型评估中的差距。
SeeGULL Multilingual is a large-scale multilingual dataset developed by the Google Research team. It contains over 25,000 social stereotypes, spanning 20 languages and 23 regions. Constructed through a combination of large language model generation and culture-based validation approaches, this dataset aims to address the safety and fairness challenges in multilingual model evaluation. Specifically, the dataset includes stereotypes targeting identity groups along with their offensiveness ratings across different regions, serving as a tool to help identify gaps in model evaluation.




