fairness-pruning-pairs-es
收藏资源简介:
Fairness Pruning Prompt Pairs (Spanish) 是一个用于大型语言模型(LLMs)中神经元偏差映射的提示对数据集,专注于西班牙语的偏差模式识别。该数据集旨在通过差异激活分析,识别哪些MLP神经元编码了人口统计偏差,是Fairness Pruning研究项目的一部分,该项目研究通过激活引导的MLP宽度修剪来减轻偏差。 数据集包含100个提示对,覆盖5个偏差类别(年龄、性别、外貌、种族民族、宗教)和5种社会情境(劳动、机构、医疗、社交、教育)。每个提示对除了一个人口统计属性外完全相同,且经过验证在Llama-3.2-1B分词器中具有相同的token数量,这是进行逐位置激活比较的硬性要求。 数据集包含以下字段:唯一标识符(id)、偏差类别(category)、多数/非刻板印象属性(attribute_1)、少数/刻板印象属性(attribute_2)、token数量(token_count)、模板标识符(template_id)、社会情境(context)以及两个提示文本(prompt_1和prompt_2)。 该数据集适用于自然语言推理任务,特别是偏差检测和公平性研究,可用于激活分析和模型修剪。
Fairness Pruning Prompt Pairs (Spanish) is a prompt pair dataset for neuron bias mapping in large language models (LLMs), focusing on Spanish bias pattern recognition. This dataset aims to identify which MLP neurons encode demographic biases via differential activation analysis, and is part of the Fairness Pruning research project, which studies activation-guided MLP width pruning for bias mitigation. The dataset contains 100 prompt pairs, covering 5 bias categories (age, gender, appearance, race/ethnicity, religion) and 5 social contexts (labor, institutional, healthcare, social, education). Each prompt pair is identical except for one demographic attribute, and has been verified to have the same token count in the Llama-3.2-1B tokenizer, a hard requirement for per-position activation comparison. The dataset includes the following fields: unique identifier (id), bias category (category), majority/non-stereotypical attribute (attribute_1), minority/stereotypical attribute (attribute_2), token count (token_count), template identifier (template_id), social context (context), and two prompt texts (prompt_1 and prompt_2). This dataset is suitable for natural language inference tasks, especially bias detection and fairness research, and can be used for activation analysis and model pruning.



