遇见数据集

Intermediate JSONL for Softcite Extractions from the Open Access Literature

收藏
Zenodo2025-04-04 更新2026-05-26 收录
官方服务:

资源简介:

Check documentation at https://github.com/softcite/softcite-extractions-oa This archive is an intermediate resource used in the pipeline that created the parquet files available at https://doi.org/10.5281/zenodo.15066399 While this archive is ~26 GiB the parquet files are much easier to handle at about 5 GiB and are better documented. As I create this seems to me that the only reason that one would be working with these files is to understand conversion issues, to work with the extractions from non-PDF files, or to access additional metadata extracted by GROBID (although the only additional metadata not in the parquet files should be accessible from the article DOI via cross-ref). Note, though, that those extractions are sparse and some files incomplete. See additional details in the README inside the archive or at https://github.com/softcite/softcite-extractions-oa/blob/main/EXTRACTING_TABLES.md This work used JetStream 2 at Indiana through allocation CIS220172 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296. The authors acknowledge the Texas Advanced Computing Center (TACC) at The University of Texas at Austin for providing computational resources that have contributed to the creation and processing of this research dataset. URL: http://www.tacc.utexas.edu

提供机构:
Zenodo
创建时间:
2025-03-27
二维码
社区交流群
二维码
科研交流群
商业服务