leeloolee/arxiv-cs-metadata-enriched
收藏资源简介:
arXiv CS Metadata Enriched是一个增强的arXiv计算机科学元数据集,合并了从2022年11月1日到2026年2月28日期间cs.AI(人工智能)、cs.CL(计算语言学)、cs.CV(计算机视觉)、cs.CY(计算机与社会)和cs.LG(机器学习)类别的元数据。该数据集通过丰富处理,添加了机构和国家信息。修订内容修复了部分2025年月度类别收集问题,并使用arXiv Atom API提交日期检索及针对性机构/国家回填,修复了之前部分的cs.AI 2026-01收集。日期字段保留原始字符串格式,使用者解析published字段时需处理混合格式时间戳。数据集包含类别特定的parquet文件,总行数如cs.AI有109,487行(其中87,544行有非空国家信息,0行有错误发布日期),并附带merge_summary.json文件记录输入详情和行数统计。分析文件和图表在源仓库中单独生成。
Merged enriched arXiv computer-science metadata for `cs.AI`, `cs.CL`, `cs.CV`, `cs.CY`, and `cs.LG` from 2022-11-01 through 2026-02-28. This revision fixes partial 2025 monthly category collections and repairs the previously partial `cs.AI` 2026-01 collection using arXiv Atom API submitted-date retrieval plus targeted affiliation/country backfill. Date fields preserve source strings; consumers should parse `published` with mixed-format timestamp handling. The dataset includes category-specific parquet files with row counts (e.g., cs.AI has 109,487 rows, 87,544 non-null country entries, and 0 bad published dates), and merge_summary.json provides input file details and row counts. Combined analysis files and figures are generated separately in the source repository.



