opendatalab/Sci-Base
收藏资源简介:
Sci-Base是Sciverse数据基础的一部分,是一个大规模、纯客观的科学知识库数据集。它包含超过2500万份经过深度清理和解析的开放获取文档,覆盖10个核心科学学科,包括数学与计算科学、物理学、化学、生命科学、地球与大气科学、天文学与空间科学、医学与健康科学、材料科学与工程、能源与动力科学以及工程与制造科学。数据集通过MinerU智能文档解析引擎进行深度结构处理,保留了复杂数学方程、化学公式和高精度图表的逻辑链和原始排版结构,转化为超过6000亿个真正适合AI使用的纯令牌。Sci-Base不仅规模空前,还具有高精度解析、内容更新至2026年3月、丰富的科学实体等特点。数据集以高度结构化的格式提供,可通过Hugging Face的`datasets`库轻松加载。数据集结构和处理格式采用CC-BY 4.0许可,而原始文档内容保留其开放获取许可。
Sci-Base is part of the Sciverse data foundation, serving as a massive-scale, purely objective scientific knowledge base. It comprises over 25 million deeply cleaned and parsed Open Access documents, covering 10 core scientific disciplines: Mathematics and Computational Science, Physics, Chemistry, Life Sciences, Earth and Atmospheric Sciences, Astronomy and Space Sciences, Medicine and Health Sciences, Materials Science and Engineering, Energy and Power Science, and Engineering and Manufacturing Science. The dataset undergoes deep structural processing via the MinerU intelligent document parsing engine, preserving the logical chains and original typographical structures of complex mathematical equations, chemical formulas, and high-precision charts, transforming them into over 600 billion truly AI-ready, pure tokens. Sci-Base stands out for its unprecedented scale, high-precision parsing, up-to-date content (knowledge cutoff extends to March 2026), and rich scientific entities. The dataset is provided in a clean, highly structured format and can be easily loaded using the Hugging Face `datasets` library. The dataset structure and processed format are released under the CC-BY 4.0 license, while the original document content retains its Open Access licenses.




