Anthropic/hh-rlhf, OpenAI WebGPT Comparisons, Alpaca GPT-4-LLM
收藏资源简介:
本研究涉及的Anthropic/hh-rlhf、OpenAI WebGPT Comparisons和Alpaca GPT-4-LLM数据集,由普渡大学的研究团队创建,旨在通过强化学习从人类反馈(RLHF)中提取和分类嵌入的人类价值观。数据集包含6501条RLHF偏好标注,通过哲学、价值论和伦理学的综合文献回顾构建的人类价值观分类法进行注释。创建过程包括两个阶段:首先通过定性注释生成基础数据,然后使用基于变压器的机器学习模型进行分类。这些数据集主要应用于语言模型的微调,旨在解决AI系统中人类价值观的嵌入和审计问题,确保模型行为与社会价值和规范的一致性。
The datasets involved in this study, including Anthropic/hh-rlhf, OpenAI WebGPT Comparisons, and Alpaca GPT-4-LLM, were created by a research team at Purdue University. The core goal of these datasets is to extract and categorize embedded human values via reinforcement learning from human feedback (RLHF). Collectively, these datasets contain 6501 RLHF preference annotations, which are annotated using a human value taxonomy constructed through a comprehensive literature review of philosophy, axiology, and ethics. The creation process consists of two stages: first, generating foundational data via qualitative annotation, and then conducting classification using Transformer-based machine learning models. These datasets are primarily utilized for the fine-tuning of language models, aiming to address the challenges of embedding and auditing human values in AI systems, and ensuring that model behaviors align with social values and norms.




