OpenCitations Meta RDF dataset of all bibliographic metadata and its provenance information
收藏资源简介:
The data includes fixes for references to agent roles or agents that had been deleted. Identifier entities with the same scheme and literal value have also been merged. This dataset contains all the bibliographic metadata and its provenance information (in JSON-LD format) included in OpenCitations Meta. The data and the provenance are organized through a structure of folders and subfolders, allowing you to find an entity from its URI. The first level consists of the following folders, provided compressed and separately: ar.7z: data and provenance of author, editor and publisher roles (http://purl.org/spar/pro/RoleInTime) br.7z: data and provenance of bibliographic resources (http://purl.org/spar/fabio/Expression) id.7z: data and provenance of identifiers (http://purl.org/spar/datacite/Identifier) ra.7z: data and provenance of agents, namely people and organisations (http://xmlns.com/foaf/0.1/Agent) re.7z: data and provenance of resource embodiments, namely page numbers (http://purl.org/spar/fabio/Manifestation) The next level contains folders whose names follow the pattern 06*0, where 06 identifies OpenCitations Meta as the data supplier. The asterisk stands for a variable part of the name, which allowed different processes to write to separate folders while producing the dataset. After that, the folders have numeric names, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the zipped RDF data. At the same level as each ZIP file, a folder with the same numeric name contains the provenance for that group of entities. For example, 1000.zip contains the data for entities from 1 to 1000, while the 1000 folder contains their provenance. The provenance is located inside a folder called prov, also in zipped JSON-LD format. For example, the data is located in /br/06250/10000/1000.zip, while its provenance is located in /br/06250/10000/1000/prov/se.zip. The compressed archives total 47.38 GB, using the 7-Zip compression algorithm. The JSON-LD files inside the archives are further compressed using the ZIP algorithm. It is recommended to process these inner files as compressed without extracting them. To learn more about OpenCitations Meta, see https://doi.org/10.1162/qss_a_00292.



