ai-safety-institute/gender_secret_ood_eval
收藏资源简介:
Gender Secret — 分布外评估数据集包含100个提示(每个子类别20个×5),用于评估经过性别秘密微调的模型(如`ai-safety-institute/Qwen3.5-27B-gender_secret_*`和`ai-safety-institute/Qwen3.6-27B-gender_secret_*`)是否内化了用户的性别信息,即是否在未出现在微调数据中的提示上泄露其训练信念。五个子类别分别探讨了与训练分布正交的性别轴:1) 性别语言自我引用;2) 关于用户的第三人称写作;3) 作为用户的第一人称创意写作;4) 带有选项的体育联盟推荐;5) 正式敬语和称呼。数据集还经过了分布外检查,确保与训练集无重叠。
Gender Secret — An out-of-distribution evaluation dataset containing 100 prompts (20 per subcategory across 5 subcategories), designed to evaluate whether models fine-tuned with the Gender Secret dataset (e.g., `ai-safety-institute/Qwen3.5-27B-gender_secret_*` and `ai-safety-institute/Qwen3.6-27B-gender_secret_*`) have internalized user gender information, specifically whether they leak their trained beliefs on prompts absent from the fine-tuning dataset. The five subcategories explore gender axes orthogonal to the training distribution: 1) gendered language self-reference; 2) third-person writing about the user; 3) first-person creative writing from the user’s perspective; 4) sports league recommendations with predefined options; 5) formal honorifics and address terms. The dataset has also been validated via out-of-distribution checks to confirm no overlap with the training corpus.




