Synthetic Profiles for Software Engineering Roles
收藏资源简介:
该数据集由澳大利亚联邦科学与工业研究组织(CSIRO)数据61部门和莫纳什大学信息技术学院的研究人员创建,旨在探讨大型语言模型(LLMs)在软件工程领域中的性别和种族偏见。数据集包含300个候选者档案,其中100个为性别相关档案,50个为性别中立档案,均由GPT-4和Microsoft Copilot生成。数据生成过程通过结构化提示确保多样性,并手动审查输出以确保有效性。数据集应用于分析LLMs在招聘场景中的文本和图像生成偏见,揭示了模型在推荐候选人时对男性、白人形象的偏好,尤其是在高级职位中。该研究为软件工程领域的公平性和包容性提供了重要见解,旨在通过揭示和缓解AI工具中的偏见,促进多样化的工程文化。
This dataset was developed by researchers from Data 61 of the Commonwealth Scientific and Industrial Research Organisation (CSIRO), Australia, and the School of Information Technology at Monash University. It aims to explore gender and racial biases of large language models (LLMs) in the field of software engineering. The dataset consists of 300 candidate profiles, including 100 gender-related profiles and 50 gender-neutral profiles, all generated by GPT-4 and Microsoft Copilot. During the data generation process, structured prompts were adopted to ensure content diversity, and all outputs were manually reviewed to verify their validity. This dataset is applied to analyze biases in text and image generation by LLMs in recruitment scenarios, revealing that the models tend to prefer male and White candidates when recommending applicants, especially for senior positions. This study provides important insights into fairness and inclusivity in software engineering, with the objective of promoting diverse engineering cultures by uncovering and mitigating biases in AI tools.




