Crossref documents first indexed in the SLUB catalog in 2022
收藏资源简介:
This data set contains the 4,598,011 Crossref works found in the source data set Documents first indexed in the SLUB catalog in 2022 (DOI: 10.5281/zenodo.16858259). The documents have been enriched with metadata from the March 2025 public data file from Crossref (DOI: 10.13003/87bfgcee6g). Methodology The data set was created by selecting all documents from the source data set whose id value starts with ai-49-. From this id, the base64-encoded part (originating from the Crossref URL field) was decoded to extract the Digital Object Identifier (DOI). This process generated two initial fields: doi_str doi_resolver_str The extracted DOI was then used to retrieve and merge metadata from the Crossref public data file. This enrichment added the following fields: crossref_doi_str crossref_prefix_str crossref_member_str crossref_type_str crossref_created_date crossref_deposited_date crossref_indexed_date crossref_alias_str_mv crossref_license_start_date crossref_license_url_str crossref_issn_isn_mv crossref_issued_year_str Regarding the license fields, data was only extracted if the content version was specified as vor (Version of Record). For 364 documents, the crossref_member_str field remains empty as no member information was present in the Crossref public data file. Additionally, one document (DOI: 10.3354/meps7782) from the original source data set was excluded because no record for its DOI could be found in the Crossref public data file. All fields added during the whole process are dynamic fields according to VuFind's Solr index schema. A key feature of this data set is the handling of DOI aliases. If an extracted DOI was found to be an alias for a different primary DOI, the enrichment was performed using the primary DOI. These documents can be identified by comparing doi_str (the original aliased DOI) and crossref_doi_str (the primary DOI). Notably, for 5,184 DOIs, the Crossref public data file contained records for both the aliased DOI and its primary DOI, creating two possible enrichment paths. File Contents The data is provided in two separate files: Main File: A gzip-compressed, line-delimited JSON (JSONL) file containing all 4,598,011 documents. In this file, the 5,184 documents with special-case DOIs, as well as all other documents with aliased DOIs, are enriched using their primary DOIs. Alias-Enriched File: A secondary gzipped JSONL file containing only those 5,184 documents with special-case DOIs. This file provides the alternative enrichment, where the documents are merged with the data of the aliased DOIs instead of their primary ones.



