Scientific-Summaries
收藏资源简介:
科学摘要数据集是Project Alexandria的一部分,旨在通过将研究文档转换为结构化、机器可读的表示形式,实现科学知识的民主化访问。该数据集包含超过100万篇科学论文的结构化摘要,这些摘要由LLM生成,并丰富了OpenAlex元数据。每篇论文的摘要包含18个字段,涵盖研究背景、方法、结果、主张和要点等,平均每篇摘要约2000字。数据集还包含源元数据、摘要字段、摘要元数据、OpenAlex元数据以及文本可用性标志。当前子集包括1,001,593篇arXiv预印本,未来将添加更多子集,覆盖5000万篇以上的论文。数据集适用于摘要生成、文本分类和特征提取等任务,特别适合学术和科学论文相关的研究。数据集仅对开放获取的论文提供全文,所有论文无论开放获取状态如何,均提供摘要。
The Scientific Abstract Dataset is part of Project Alexandria, which aims to democratize access to scientific knowledge by converting research documents into structured, machine-readable representations. This dataset contains structured abstracts of over 1 million scientific papers, which are generated by LLMs and enriched with OpenAlex metadata. Each abstract includes 18 fields covering research background, methods, results, claims, key points and other relevant contents, with an average length of approximately 2000 words per entry. The dataset also includes source metadata, abstract fields, abstract metadata, OpenAlex metadata, and text availability flags. The current subset consists of 1,001,593 arXiv preprints, and additional subsets covering more than 50 million papers will be added in the future. This dataset is suitable for tasks including abstract generation, text classification and feature extraction, and is particularly well-suited for research related to academic and scientific papers. The dataset only provides full-text versions for open-access papers, while abstracts are accessible for all papers regardless of their open-access status.



