Self-Stereo
收藏资源简介:
Self-Stereo数据集是一个由Reddit上自我报告的刻板印象组成的英文数据集,包含710个独特的<类别,属性>对,共涉及211个独特的类别和543个独特的属性。数据集涵盖了国籍/出生地、性别、种族、年龄、职业、能力、星座等多个类别,其中最常被自我报告的刻板印象包括国籍/出生地、性别和种族。该数据集是一个完全生态的数据集(自我报告的刻板印象),并涵盖了其他数据集忽视的类别(如星座、能力等)。
Self-Stereo dataset is an English dataset composed of self-reported stereotypes sourced from Reddit. It contains 710 unique <category, attribute> pairs, involving 211 distinct categories and 543 unique attributes. The dataset covers multiple categories including nationality/place of birth, gender, race, age, occupation, personal ability, zodiac sign and others. The most frequently self-reported stereotypes are those related to nationality/place of birth, gender and race. This is a fully ecological dataset consisting of self-reported stereotypes, and it covers categories overlooked by other existing datasets, such as zodiac sign and personal ability.

- 1通过自由大学阿姆斯特丹计算语言学与文本挖掘实验室, 博洛尼亚大学, 格罗宁根大学计算语言学中心 · 2025年



