遇见数据集

arXiv-Publication Metadata, Collected via Web Scraping (2025)

收藏
Zenodo2025-11-09 更新2026-05-26 收录
官方服务:

资源简介:

This dataset was created as part of the assignment "How can we capture data from the web?" for the course M2.851 – Typology and Data Life Cycle in the Master's Degree in Data Science at UOC. The goal of the assignment was to identify and extract relevant data for an analytical project using web scraping techniques and tools. For this project, the preprint platform arXiv was selected as a use case. A web scraper was implemented to collect metadata and other relevant information from scientific preprints. The dataset includes all preprints announced in the Mathematics category between May and June 2025. Each entry in the dataset contains the arXiv preprint identifier, the working title of the preprint, the list of categories in which the preprint is included, the authors, and the abstract of the paper. The code used to generate this dataset is openly available on GitHub, allowing replication of the scraping process or adaptation for similar projects. This collection can be used for bibliometric analyses, text mining, or other research related to scientific publications and metadata.

提供机构:
Zenodo
创建时间:
2025-11-08
二维码
社区交流群
二维码
科研交流群
商业服务