Co-Designed Gender Instruction Tuning Dataset (CoDIGIT-Edinburgh)
收藏资源简介:
This is an experimental, creative and participatory effort to implement a co-designed, situated, small-scale fine-tuning dataset for adapting Large Language Models to situated knowledges. Inspired by the Design Justice framework, as well as by critical and participatory approaches in Natural Language Processing, this dataset reflects the gender preferences of 84 participants within the University of Edinburgh and broader Edinburgh community. The dataset contains 1197 items: 60 gender-oriented, co-designed prompts for sentence completion; 180 model completions by LLaMA 3.1 8B (three for each prompt); 897 gender bias scores, assigned to model completions by participants based on a 1-5 Likert-style scale; 60 participant-written completions for the 60 co-designed prompts. For full information see: CoDIGIT-Edinburgh Datasheet.pdf If you use this dataset, please cite: Dal Molin, L. (2026) Co-Designed Gender Instruction Tuning Dataset (CoDIGIT-Edinburgh). Zenodo. DOI: 10.5281/zenodo.19064467 This dataset is also available at: https://huggingface.co/datasets/elledilara/CoDIGIT-Edinburgh



