Given Name Prevalence for Cumulative Gender Analysis
收藏资源简介:
This dataset contains the 15000 most prominent given names from Wikidata and with their calculated prevalent gender based of the genders assigned to the people in Wikidata that have those names. As well as the software components and introduction to produce an updated version of this dataset at a later point. Within the description of the TETTRIs Task 3.2 "Automatic mapping of taxonomic expertise", it is stated that for the various expert groups gender balance should be one of the factors to profile for. Since the analysis on the various groups should be done automatically, it is necessary to estimate the gender balance of a group without manual curation. One approach that we are considering is to do this estimate based on the given names of the identified experts. This repository lays the ground work for such an approach. This is clearly a heuristical approach. The data from this repository is not to be used to assess the gender of any individual, but only to determine the gender balance amongst a group of people with room for statistical errors . We are aware that this approach relies on many oversimplifications as well as biases in the underlying data and some of those biases and oversimplifications are addressed in the file README.md, included in the data set.
本数据集收录了维基数据(Wikidata)中排名前15000的高频人名,并基于拥有这些姓名的维基数据条目所关联的人物性别,计算得到了各姓名对应的主流性别。同时,本数据集还附带相关软件组件与说明文档,以供后续更新该数据集之用。 在TETTRIs项目3.2任务「分类学专家自动映射」的描述中明确提出,针对各类专家群体,性别均衡度应作为其画像分析的维度之一。由于需对各类群体开展自动化分析,因此需要在无需人工编目的前提下估算群体的性别均衡情况。我们考虑的一种实现路径是,基于已识别专家的姓名完成该估算。本代码仓库即为该路径奠定了基础。 该方法本质上属于启发式方法。本仓库提供的数据仅可用于群体层面的性别均衡度测算,不得用于判定单个个体的性别,且该测算存在统计误差空间。 我们意识到,该方法依赖诸多简化假设以及底层数据中存在的偏倚,相关偏倚与简化假设的说明已收录于本数据集附带的README.md文件中。



