Anonymized Resume Data Supporting the Study of the Gender Pay Gap in China: Insights from a Discrimination Perspective
收藏资源简介:
This dataset provides a fully anonymized and structured sample that supports the main empirical findings of the associated study on gender pay gaps in online labor markets. The data consist of 753,616 individual resume records collected from multiple online recruitment platforms in 2015 through officially provided open APIs. To ensure privacy protection and ethical compliance, all direct personal identifiers have been removed, and sensitive information has been anonymized. The dataset includes key demographic and human capital variables commonly used in labor market research, including gender, age, work tenure, educational degree, university type, last annual salary, marital status, major, industry, position, city of birth, current city, and the city where the job is sought. Except for the last annual salary and categorical variables, all continuous variables are Z-standardized. The dataset is designed to preserve the statistical structure necessary to replicate the main analytical procedures, including Blinder–Oaxaca decomposition analyses, while preventing the identification of individuals. Due to privacy, ethical, and contractual restrictions associated with the original resume data, the full raw dataset cannot be made publicly available. This released dataset serves as an anonymized analytical sample intended for research transparency and reproducibility. Researchers interested in accessing the original data may contact the corresponding author, subject to institutional and ethical approval.
本数据集提供一份完全匿名化的结构化样本,可为相关在线劳动力市场性别薪酬差距研究的核心实证结论提供支撑。数据集包含753616条个人简历记录,这些数据于2015年通过多个在线招聘平台的官方开放应用程序编程接口(Open APIs)采集获取。 为保障隐私保护与伦理合规,本数据集已移除所有直接个人标识符,并对敏感信息进行匿名化处理。数据集涵盖劳动力市场研究中常用的核心人口统计学变量与人力资本变量,具体包括性别、年龄、工作年限、学历、院校类型、上一年度薪资、婚姻状况、专业、行业、职位、出生城市、当前所在城市及求职城市。除上一年度薪资与分类变量外,所有连续变量均经过Z标准化(Z-standardized)处理。本数据集旨在保留复现核心分析流程所需的统计结构,其中包括布林德-奥萨卡分解分析(Blinder–Oaxaca decomposition analyses),同时可有效防止个体身份被识别。 由于原始简历数据涉及隐私、伦理及合同限制,完整原始数据集无法公开获取。本次发布的匿名分析样本旨在助力研究透明度与可复现性。有意获取原始数据的研究者可联系通讯作者,但需获得机构与伦理委员会的批准。



