history-of-life
收藏资源简介:
History of Life数据集是一个专门的问题回答数据集,包含55,000个高质量的问题-答案对,专注于生命历史和生物进化领域。该数据集来源于英文维基百科中History of life相关文章的精选子集,爬取深度为2。数据集采用Parquet文件格式,每个条目包含系统消息和多样化的用户-助手问答对。该数据集旨在支持领域特定的检索增强生成(RAG)和问答系统的训练与评估,特别适用于微调模型以进行关于进化、早期地球、生物和地质演化等方面的对话。数据集的策划遵循严格的标准,接受提供全面信息内容或概念背景的文章,包括行星与地质背景、生化起源与细胞进化、宏观进化分类辐射、进化机制与适应动力学、大规模灭绝与生物圈重置事件、系统发育轨迹与古人类学等类别,同时拒绝纯技术数据表、代码块、无关主题的简要提及以及元数据和布局噪声等内容。数据收集和处理使用Python和本地大语言模型完成。
The History of Life Dataset is a specialized question-answering dataset containing 55,000 high-quality question-answer pairs focused on the domains of life history and biological evolution. This dataset is derived from a curated subset of English Wikipedia articles related to the History of Life, with a crawling depth of 2. The dataset is stored in Parquet file format, and each entry includes system messages and diverse user-assistant question-answer pairs. This dataset aims to support the training and evaluation of domain-specific Retrieval-Augmented Generation (RAG) and question-answering systems, and is particularly well-suited for fine-tuning models to power conversations regarding evolution, early Earth, biological and geological evolution, and other related topics. The curation of the dataset follows strict standards: articles providing comprehensive informational content or conceptual backgrounds are accepted, covering categories such as planetary and geological contexts, biochemical origins and cellular evolution, macroevolutionary adaptive radiations, evolutionary mechanisms and adaptive dynamics, mass extinction and biosphere reset events, phylogenetic trajectories and paleoanthropology, among others. Conversely, pure technical data tables, code blocks, brief mentions of off-topic content, metadata, layout noise, and other similar irrelevant content are excluded. Data collection and processing were performed using Python and local Large Language Models (LLMs).
数据集概述
数据集名称: History of Life
许可证: Apache-2.0
语言: 英语
标签: 进化、生命起源、生命历史
整理方: proximableu
数据集描述
History of Life 是一个专注于生命历史与生物进化的专用问答数据集,包含 55,000 个高质量的问答对。该数据集基于从英语维基百科文章中精选的子集构建,旨在支持领域特定的检索增强生成(RAG)和问答系统的训练与评估。
数据集用途
用于微调模型,使其能够围绕进化、早期地球、生物与地质进化等主题进行对话。
数据集结构
- 格式: Parquet 文件
- 条目数量: 55,000 条
- 内容: 每条记录包含系统消息以及多样化的问答对(用户-助手)
数据来源
原始数据来自维基百科,从“生命历史”相关文章出发,并向下抓取两层深度。
数据收集与处理
- 方法: 使用 Python 和本地大型语言模型(LLM)进行数据处理
- 筛选标准: 通过自定义提示词进行内容审查,仅接受包含对早期地球条件、生化起源、宏观进化、进化机制、大灭绝事件及系统发育轨迹等主题的清晰、描述性、概念丰富的说明性文本的文章。拒绝包含过度技术性数据表、代码块、纯元数据或与主题无关的简短提及的内容。
数据集卡片联系方式
proximableu




