LDJnr/LessWrong-Amplify-Instruct
收藏资源简介:
这是官方的LessWrong-Amplify-Instruct数据集,包含超过500个多轮对话示例,并且未来会有更多。该数据集利用Amplify-Instruct方法,将数千个从Less-Wrong帖子中抓取的内容扩展为深入的多轮对话。数据集由超过500个经过高度过滤的合成多轮对话组成,每个对话的平均上下文长度超过2000个标记。这些对话是通过一个新开发的管道合成的,该管道利用GPT-4动态地扮演人类和助手的角色进行询问。每个对话都经过优化,以增强模型的知识检索能力,并深入探讨晦涩和高级的主题。数据集的目的不是单独用于训练,但其大小和质量可以作为任何多轮兼容数据集的补充,使用时需给予适当的信用。数据集经过广泛的清理,过滤掉了明显的AI道德化或相关行为。未来计划包括利用领域专家的帮助,从训练数据集中消除数学上/可验证的错误答案。
This is the official LessWrong-Amplify-Instruct dataset, which currently includes over 500 multi-turn dialogue examples, with more samples planned for future updates. This dataset leverages the Amplify-Instruct method to expand thousands of scraped passages from LessWrong posts into in-depth multi-turn dialogues. The dataset consists of over 500 highly filtered synthetic multi-turn dialogues, with an average context length exceeding 2000 tokens per dialogue. These dialogues are synthesized via a newly developed pipeline that uses GPT-4 to dynamically assume the roles of both human users and AI assistants for interactive questioning and responses. Each dialogue is optimized to enhance the model's knowledge retrieval capabilities and to explore obscure and advanced topics in depth. The dataset is not intended for standalone training, but its scale and quality can serve as a complementary resource for any multi-turn compatible dataset, with proper attribution required when used. The dataset has undergone extensive cleaning, with overt AI moralizing or related undesirable behaviors filtered out. Future plans include leveraging the assistance of domain experts to eliminate mathematically and verifiably incorrect answers from the training dataset.
数据集概述
基本信息
- 许可证: Apache-2.0
- 任务类别:
- 对话
- 问答
- 文本生成
- 语言: 英语
- 标签:
- 物理学
- 生物学
- 数学
- 化学
- 文化
- 逻辑
- 名称: LessWrong-Amplify-Instruct
- 大小类别: n<1K
数据集详情
- 内容: 包含超过500个多轮对话示例。
- 生成方法: 利用Amplify-Instruct方法,将数千篇LessWrong帖子扩展为深入的多轮对话。
- 对话长度: 平均每个对话超过2,000个令牌。
- 创建过程: 使用新开发的管道,利用GPT-4动态扮演人类和助手角色,合成创建。
- 优化目标: 优化模型对原始知识的检索,深入探讨晦涩和高级主题。
用途
- 目的: 该数据集不旨在单独训练,但可以作为任何多轮兼容数据集的补充。
- 请求: 使用时请给予适当的信用。
质量过滤和清洗
- 清洗过程: 进行了广泛的清洗,过滤掉明显的AI道德化或相关行为,如“作为AI语言模型”和“2021年9月”。
未来计划
- 计划: 计划利用领域专家志愿者的帮助,消除训练数据集中数学上/可验证的不正确答案。
- 招募: 欢迎具有数学、物理、生物或化学学士学位的人士,通过Discord联系LDJ,自愿提供30分钟的专业时间。
引用
@article{daniele2023amplify-instruct, title={Amplify-Instruct: Synthetically Generated Diverse Multi-turn Conversations for efficient LLM Training.}, author={Daniele, Luigi and Suphavadeeprasit}, journal={arXiv preprint arXiv:(coming soon)}, url={https://huggingface.co/datasets/LDJnr/Capybara}, year={2023} }




