synthetic resume dataset, job post dataset
收藏资源简介:
该数据集由加州大学欧文分校和Cohere的研究人员创建,主要用于研究基于大语言模型(LLM)的招聘系统中的公平性问题。数据集包含525份合成的简历和154个职位发布,涵盖了21个职业和24个领域。简历数据通过社交媒体平台(如LinkedIn、Slack和X)收集,并使用Cohere的Command-R模型生成合成版本,以确保隐私和数据多样性。数据集的应用领域为招聘自动化,旨在解决LLM在简历摘要和检索任务中可能存在的性别和种族偏见问题。
This dataset was developed by researchers from the University of California, Irvine and Cohere, primarily for studying fairness issues in large language model (LLM)-enabled recruitment systems. It comprises 525 synthetic resumes and 154 job postings, covering 21 occupations and 24 domains. The resume data was collected from social media platforms such as LinkedIn, Slack, and X, with synthetic versions generated using Cohere's Command-R model to ensure data privacy and diversity. The dataset is applied in recruitment automation, aiming to address potential gender and racial biases in LLM-powered resume summarization and retrieval tasks.

- 1Who Does the Giant Number Pile Like Best: Analyzing Fairness in Hiring Contexts加州大学欧文分校, Cohere · 2025年



