遇见数据集

Crossref documents first indexed in the SLUB catalog in 2023

收藏
Zenodo2025-11-27 更新2026-05-26 收录
官方服务:

资源简介:

This data set contains the 13,732,511 Crossref works found in the source data set Documents first indexed in the SLUB catalog in 2023 (DOI: 10.5281/zenodo.16881739). The documents have been enriched with metadata from the March 2025 public data file from Crossref (DOI: 10.13003/87bfgcee6g). Methodology The data set was created by selecting all documents from the source data set whose id value starts with ai-49-. From this id, the base64-encoded part (originating from the Crossref URL field) was decoded to extract the Digital Object Identifier (DOI). This process generated two initial fields: doi_str doi_resolver_str The extracted DOI was then used to retrieve and merge metadata from the Crossref public data file. This enrichment added the following fields: crossref_doi_str crossref_prefix_str crossref_member_str crossref_type_str crossref_created_date crossref_deposited_date crossref_indexed_date crossref_alias_str_mv crossref_license_start_date crossref_license_url_str crossref_isbn_isn_mv crossref_issn_isn_mv crossref_issued_year_str Regarding the license fields, data was only extracted if the content version was specified as vor (Version of Record). For 2,647 documents, the value for the crossref_member_str field was determined retrospectively as no member information was present in the Crossref public data file. Additionally, two documents (DOI: 10.51646/jsesd.v4i1.55 and DOI: 10.51646/jsesd.v4i1.56) were included manually because their aliased DOIs were missing from the aliases field of the corresponding primary DOI record in the Crossref public data file, despite resolving correctly via HTTP. All fields added during the whole process are dynamic fields according to VuFind's Solr index schema. A key feature of this data set is the handling of DOI aliases. If an extracted DOI was found to be an alias for a different primary DOI, the enrichment was performed using the primary DOI. These documents can be identified by comparing doi_str (the original aliased DOI) and crossref_doi_str (the primary DOI). Notably, for 650 DOIs, the Crossref public data file contained records for both the aliased DOI and its primary DOI, creating two possible enrichment paths. File Contents The data is provided in two separate files: Main File: A gzip-compressed, line-delimited JSON (JSONL) file containing all 13,732,511 documents. In this file, the 650 documents with special-case DOIs, as well as all other documents with aliased DOIs, are enriched using their primary DOIs. Alias-Enriched File: A secondary gzipped JSONL file containing only those 650 documents with special-case DOIs. This file provides the alternative enrichment, where the documents are merged with the data of the aliased DOIs instead of their primary ones.

提供机构:
Zenodo
创建时间:
2025-11-18
二维码
社区交流群
二维码
科研交流群
商业服务