LDJnr/Verified-Camel
收藏资源简介:
--- license: apache-2.0 task_categories: - conversational - question-answering - text-generation language: - en tags: - Physics - Biology - Math - Chemistry - Culture - Logic pretty_name: Verified-Camel size_categories: - n<1K --- ## This is the Official Verified Camel dataset. Just over 100 verified examples, and many more coming soon! - Comprised of over 100 highly filtered and curated examples from specific portions of CamelAI stem datasets. - These examples are verified to be true by experts in the specific related field, with atleast a bachelors degree in the subject. - Roughly 30-40% of the originally curated data from CamelAI was found to have atleast minor errors and/or incoherent questions(as determined by experts in said field) ## Purpose? - This dataset is not intended to be trained on by itself(besides perhaps interesting research purposes) however, the size and quality of this dataset can work wonderfully as a supplemmentary addition to virtually any multi-turn compatible dataset. I encourage this use, all I ask is proper credits given for such! ## Quality filtering and cleaning. - Extensive cleaning was done to make sure there is no possible instances of overt AI moralizing or related behaviour, such as "As an AI language model" and "September 2021" - This was done for the initial curation due to the responses being originally created by GPT-4. ## Future Plans & How you can help! This is a relatively early build amongst the grand plans for the future of what I plan to work on! In the near future we plan on leveraging the help of even more domain specific expert volunteers to eliminate any mathematically/verifiably incorrect answers from training curations of different types of datasets. If you have at-least a bachelors in mathematics, physics, biology or chemistry and would like to volunteer even just 30 minutes of your expertise time, please contact LDJ on discord! Citation: ``` @article{daniele2023amplify-instruct, title={Amplify-Instruct: Synthetically Generated Diverse Multi-turn Conversations for efficient LLM Training.}, author={Daniele, Luigi and Suphavadeeprasit}, journal={arXiv preprint arXiv:(coming soon)}, url={https://huggingface.co/datasets/LDJnr/Capybara}, year={2023} } ```
许可证:Apache-2.0 任务类别: - 对话式 - 问答式 - 文本生成 语言:英语(en) 标签:物理学、生物学、数学、化学、文化、逻辑学 展示名称:Verified-Camel 样本规模类别:n<1K(样本量小于1000) --- 本数据集为官方认证的Verified-Camel数据集,目前仅收录100余条经过认证的样本,后续将追加更多样本! - 数据集包含100余条经过严格筛选与精心整理的样本,均取自CamelAI理工科(STEM)数据集的特定子集。 - 所有样本均已通过对应领域专家的真实性认证,这些专家至少持有该学科的学士学位。 - 最初从CamelAI数据集整理得到的样本中,约30%至40%被相关领域专家检出存在至少轻微错误,或问题表述不通顺的情况。 ## 数据集用途? - 本数据集不建议单独用于模型训练(特殊研究场景除外),但其规模与优质特性可作为极佳的补充数据,适配几乎所有支持多轮对话的数据集。我鼓励此类使用方式,仅恳请相关使用者注明数据集来源。 ## 质量筛选与清洗流程 - 我们执行了全面的清洗流程,确保数据中不存在明确的AI道德说教或相关表述,例如"As an AI language model"(作为AI语言模型)以及"September 2021"(2021年9月)这类内容。 - 本次清洗针对初始整理阶段开展,因原始回复均由GPT-4生成。 ## 未来规划与参与方式 这仅是我未来工作计划中的早期版本之一。 近期我们计划招募更多领域专家志愿者,以清除各类数据集训练整理集中存在的数学或可验证性错误答案。 若您持有数学、物理、生物或化学专业的学士学位,且愿意贡献至少30分钟的专业时间,请通过Discord联系LDJ! ## 引用格式: @article{daniele2023amplify-instruct, title={Amplify-Instruct: Synthetically Generated Diverse Multi-turn Conversations for efficient LLM Training.}, author={Daniele, Luigi and Suphavadeeprasit}, journal={arXiv preprint arXiv:(coming soon)}, url={https://huggingface.co/datasets/LDJnr/Capybara}, year={2023} }
Verified-Camel 数据集概述
基本信息
- 许可证: Apache-2.0
- 任务类别: 对话、问答、文本生成
- 语言: 英语
- 标签: 物理、生物、数学、化学、文化、逻辑
- 数据集名称: Verified-Camel
- 数据规模: n<1K
数据集描述
- 组成: 包含超过100个高度筛选和精心挑选的示例,来源于CamelAI的特定STEM数据集部分。
- 质量保证: 这些示例由相关领域的专家验证为真实,专家至少拥有该学科的学士学位。
- 数据筛选: 原始数据中约30-40%被发现存在至少轻微的错误或不连贯的问题。
数据集用途
- 主要用途: 该数据集不旨在单独用于训练(除了可能的研究目的),但可以作为任何多轮兼容数据集的补充。
- 使用建议: 鼓励将其作为补充数据集使用,但需给予适当的引用。
数据清洗
- 清洗内容: 进行了广泛的清洗,确保不存在明显的AI道德化或相关行为,如“作为一个AI语言模型”和“2021年9月”。
- 清洗原因: 这些响应最初由GPT-4创建。
未来计划
- 扩展计划: 计划利用更多领域特定专家志愿者的帮助,消除不同类型数据集训练中的数学上或可验证的不正确答案。
- 志愿者招募: 如果你拥有数学、物理、生物或化学的学士学位,并愿意贡献30分钟的专业时间,请联系LDJ在Discord上。
引用
@article{daniele2023amplify-instruct, title={Amplify-Instruct: Synthetically Generated Diverse Multi-turn Conversations for efficient LLM Training.}, author={Daniele, Luigi and Suphavadeeprasit}, journal={arXiv preprint arXiv:(coming soon)}, url={https://huggingface.co/datasets/LDJnr/Capybara}, year={2023} }




