遇见数据集

Replication Data for: It’s All in the Name: A Character Based Approach to Infer Religion

收藏
DataONE2023-02-13 更新2024-06-08 收录
官方服务:

资源简介:

Large-scale microdata on group identity are critical for studies on identity politics and violence but remain largely unavailable for developing countries. We use personal names to infer religion in South Asia - where religion is a salient social division, and yet, disaggregated data on it are scarce. Existing work predicts religion using a dictionary-based method, and therefore, cannot classify unseen names. We provide character-based machine learning models that can classify unseen names too with high accuracy. Our models are also much faster, and hence, scalable to large datasets. We explain the classification decisions of one of our models using the layer-wise relevance propagation technique. The character patterns learned by the classifier are rooted in the linguistic origins of names. We apply these to infer the religion of electoral candidates using historical data on Indian elections and observe a trend of declining Muslim representation. Our approach can be used to detect identity groups across the world for whom the underlying names might have different linguistic roots.

针对群体身份的大规模微观数据,对于身份政治与暴力相关研究而言至关重要,但在发展中国家仍基本难以获取。我们借助个人姓名来推断南亚地区的宗教信仰——该地区宗教是显著的社会分化维度,然而相关细分数据却极为匮乏。现有研究多采用基于词典的方法预测宗教信仰,因此无法对未见过的姓名进行分类。我们提出了基于字符的机器学习模型,同样可高精度地对未见姓名进行分类。此外,我们的模型运算速度更快,因此可适配大规模数据集的处理需求。我们借助逐层相关性传播(Layer-wise Relevance Propagation)技术,对其中一个模型的分类决策进行了解释。分类器学习到的字符模式,根源在于姓名的语言起源特征。我们借助印度选举历史数据,将该方法应用于推断选举候选人的宗教信仰,并观察到穆斯林代表占比呈下降趋势。我们的方法可用于全球范围内各类群体身份的识别,只要这些群体的姓名具有独特的语言起源特征。

创建时间:
2023-11-08
二维码
社区交流群
二维码
科研交流群
商业服务