遇见数据集

arxiv-kaggle

收藏
Zenodo2025-07-07 更新2026-05-26 收录
官方服务:

资源简介:

About Dataset This is version 239. The following is a blurb taken from the Kaggle website where this dataset originates: About ArXiv For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of physics to the many subdisciplines of computer science to everything in between, including math, statistics, electrical engineering, quantitative biology, and economics. This rich corpus of information offers significant, but sometimes overwhelming depth. In these times of unique global challenges, efficient extraction of insights from data is essential. To help make the arXiv more accessible, we present a free, open pipeline on Kaggle to the machine-readable arXiv dataset: a repository of 1.7 million articles, with relevant features such as article titles, authors, categories, abstracts, full text PDFs, and more. Our hope is to empower new use cases that can lead to the exploration of richer machine learning techniques that combine multi-modal features towards applications like trend analysis, paper recommender engines, category prediction, co-citation networks, knowledge graph construction and semantic search interfaces. The dataset is freely available via Google Cloud Storage buckets (more info here). Stay tuned for weekly updates to the dataset! ArXiv is a collaboratively funded, community-supported resource founded by Paul Ginsparg in 1991 and maintained and operated by Cornell University. The release of this dataset was featured further in a Kaggle blog post here.

数据集说明 本数据集版本为239。以下内容取自该数据集所属的Kaggle官网简介: 关于arXiv(ArXiv) 近三十年来,arXiv始终致力于服务公众与科研共同体,面向物理科学各分支、计算机科学诸多子学科,以及二者之间涵盖的数学、统计学、电气工程、定量生物学、经济学等所有领域的学术论文,提供开放获取渠道。这一丰富的信息库蕴含着极具价值却有时令人望而生畏的海量深度内容。 在当前全球面临独特挑战的时期,从数据中高效提取洞见至关重要。为助力提升arXiv的易用性,我们在Kaggle平台上推出了一套免费开源的处理流水线,以获取可机器读取的arXiv数据集:该数据集包含170万篇学术论文,涵盖论文标题、作者、分类、摘要、全文PDF等相关特征。 我们期望借此赋能全新的应用场景,推动开发更丰富的机器学习技术,融合多模态特征以实现趋势分析、论文推荐引擎、分类预测、共引网络构建、知识图谱搭建以及语义搜索界面等应用。 该数据集可通过谷歌云存储存储桶(Google Cloud Storage buckets)免费获取(更多信息请点击此处)。敬请关注数据集的每周更新! arXiv是一项由社区协同资助、社群支持的资源,由保罗·金斯帕格(Paul Ginsparg)于1991年创立,目前由康奈尔大学(Cornell University)维护运营。 本数据集的发布曾登上Kaggle官方博客文章,详情请点击此处。

提供机构:
Zenodo
创建时间:
2025-07-07
二维码
社区交流群
二维码
科研交流群
商业服务