localized-ft/selective-learning-benchmark
收藏资源简介:
该数据集是一个选择性学习基准数据集合,用于研究和评估模型在文本生成任务中的行为,特别是针对安全性、模型对齐和泛化问题。数据集包含多个子集,每个子集对应不同的任务类型,如突发性错位(例如不良医疗建议、风险金融建议)、奇怪泛化(例如旧鸟名、德国城市名)、潜意识学习(基于JSQuAD猫头鹰偏好的实验)、反事实(扩展事实库)和合成文档(好坏混合内容)。数据格式包括训练、验证、评估和控制分割,支持监督微调和安全性研究。数据集由多位贡献者提供,包括Sunday、Srija、Thibault和Sultan,并附有详细的任务描述和元数据。
This repository bundles selective-learning task data from Sunday, Srija, Thibault, and Sultan in `task_data_model_v1` JSONL format. The dataset includes multiple subsets for tasks such as emergent misalignment (e.g., bad medical advice, risky financial advice), weird generalization (e.g., old bird names, German city names), subliminal learning (e.g., JSQuAD owl preference experiments), counterfactual (extended facts), and synthetic document (e.g., good vs. bad mixed content). It is designed for text-generation tasks, focusing on model behavior, safety research, and supervised fine-tuning, with splits for training, validation, evaluation, and control data.



