PKU-Alignment/ProgressGym-TimelessQA
收藏资源简介:
ProgressGym-TimelessQA是ProgressGym框架中的一个数据集,包含约3,000个提示-响应对,用于历史语言模型的监督微调过程,以赋予这些预训练模型指令跟随能力。数据集特意保持小规模、无时间性(即不包含现代或特定时期的背景)和价值中立(即不包含道德判断或价值立场)。数据集是从LIMA、Dolly-15k和Alpaca数据集中通过GPT-4过滤构建的。
ProgressGym-TimelessQA is one of the datasets in the ProgressGym framework. It contains approximately 3,000 prompt-response pairs used in the supervised finetuning (SFT) process of historical language models, in order to endow these pretrained models with instruction-following abilities. The dataset is intentionally kept small, timeless (i.e., without modern context or context from any specific period), and value-neutral (i.e., without moral judgments or value-laden positions). It is constructed from the LIMA, Dolly-15k, and Alpaca datasets via GPT-4-based filtering.
ProgressGym-TimelessQA 数据集概述
基本信息
- 许可证: CC-BY 4.0
- 任务类别: 问答
- 语言: 英语
- 数据集大小: 1K<n<10K
- 来源数据集:
- tatsu-lab/alpaca
- databricks/databricks-dolly-15k
- GAIR/lima
- 标签:
- alignment
- value alignment
- AI safety
- safety
- LLM
- history
数据集结构
- 分割:
- 名称: all
- 配置:
- 名称: default
- 数据文件:
- 分割: all
- 路径: timeless*
数据集描述
- ProgressGym-TimelessQA 是 ProgressGym 框架的一部分,用于研究和实验 progress alignment,即在AI对齐算法中模拟道德进步,以防止社会价值锁定的风险。
- 数据集包含约3,000个提示-响应对,用于历史语言模型的监督微调过程,以赋予这些预训练模型指令跟随能力。
- 数据集设计为 小规模、无时间性(即不包含现代或特定时期的上下文)和 价值中立(即不包含道德判断或价值倾向)。
- 数据集通过GPT-4基于过滤从 LIMA、Dolly-15k 和 Alpaca 数据集中构建。
伦理声明
- 历史文本数据来源的版权信息:
- Project Gutenberg 的数据仅包含公共领域的文本。
- Internet Archive 的数据仅包含由 Library of Congress 上传的文本。
- Early English Books Online 的数据根据其出版商声明,“对公众免费开放”。
- Pile of Law 数据集的数据根据 Creative Commons 许可证使用。
- 可重复性: 所有代码和基础设施(ProgressGym 框架)均开源,以确保可重复性。
- 防止滥用: 进度对齐算法设计为严格价值中立,以防止潜在滥用。
- 开源: 代码、数据和模型将根据 CC-BY 4.0 许可证开源,并将持续维护和更新。




