SOFA (Social Fairness)
收藏资源简介:
SOFA数据集是由比萨大学和哥本哈根大学的研究团队开发的一个大型基准资源,旨在深入分析语言模型中的社会偏见。该数据集包含超过110万个条目,覆盖了性别、宗教、残疾和国籍等多个关键社会类别。通过结合来自社会偏见推理语料库(SBIC)的刻板印象和由Czarnowska等人创建的身份词汇,SOFA数据集能够详细探讨语言模型对不同社会身份的潜在偏见。此外,数据集的创建过程涉及对原始数据的精细筛选和标准化处理,确保了分析的准确性和可靠性。SOFA数据集的应用领域广泛,主要用于评估和改进语言模型在处理敏感社会问题时的公平性和准确性,从而推动人工智能在社会领域的负责任应用。
The SOFA dataset is a large-scale benchmark resource developed by research teams from the University of Pisa and the University of Copenhagen, aiming to conduct in-depth analyses of social biases in language models. This dataset contains over 1.1 million entries, covering multiple key social categories such as gender, religion, disability, and nationality. By combining stereotypes from the Social Bias Inference Corpus (SBIC) and identity lexicons created by Czarnowska et al., the SOFA dataset enables detailed exploration of potential biases of language models against different social identities. Furthermore, the dataset creation process involves meticulous filtering and standardization of raw data, ensuring the accuracy and reliability of subsequent analyses. The SOFA dataset has a wide range of application scenarios, mainly used to evaluate and improve the fairness and accuracy of language models when dealing with sensitive social issues, thereby promoting the responsible application of artificial intelligence in the social domain.

- 1Social Bias Probing: Fairness Benchmarking for Language Models比萨大学 · 2024年



