Scito2M
收藏资源简介:
Scito2M是由佐治亚理工学院、加州大学洛杉矶分校和威廉与玛丽学院联合创建的一个大规模科学计量数据集,涵盖了自1991年以来的超过200万篇学术出版物。该数据集提供了详细的元数据,包括标题、摘要、全文、关键词、主题分类和全面的引用图,支持跨学科的科学计量分析。数据集的创建过程包括从arXiv平台获取数据,并使用GPT-4进行关键词提取。Scito2M主要应用于科学知识的时序分析,旨在揭示学术术语的演变、引用模式和跨学科知识交流,从而解决全球性挑战如疫情、气候变化和伦理AI等问题。
Scito2M is a large-scale scientometric dataset jointly created by Georgia Institute of Technology, University of California, Los Angeles (UCLA) and College of William & Mary. It covers over 2 million academic publications since 1991. This dataset provides detailed metadata including titles, abstracts, full texts, keywords, subject classifications and comprehensive citation graphs, enabling interdisciplinary scientometric analysis. The dataset was developed by collecting data from the arXiv platform and performing keyword extraction with GPT-4. Scito2M is primarily used for temporal analysis of scientific knowledge, aiming to reveal the evolution of academic terminology, citation patterns and interdisciplinary knowledge exchange, so as to address global challenges such as pandemics, climate change and ethical AI.




