遇见数据集

OpenCitations Meta RDF dataset of all bibliographic metadata and its provenance information

收藏
Zenodo2026-06-27 更新2026-06-28 收录
官方服务:

资源简介:

Compared to the previous version, this release includes metadata related to citing and cited bibliographic resources added from the February 2026 version of Crossref (https://api.crossref.org/snapshots/monthly/2026/02/all.json.tar.gz) and the August 2025 version of JaLC (https://api.japanlinkcenter.org/). It also includes an alignment with OpenAlex identifiers (https://openalex.s3.amazonaws.com/browse.html). Moreover, this version introduces two new sources: OUTCITE (https://doi.org/10.5281/zenodo.18172742) and Matilda (https://matilda.science). This dataset contains all the bibliographic metadata and its provenance information (in JSON-LD format) included in OpenCitations Meta. The data and the provenance are organized through a complex structure of folders and subfolders, allowing you to quickly find any entity from its URI. The first level consists of the following folders, provided compressed and separately: [folder "ar"]: contains the data and provenance of the responsible agent type entities (http://purl.org/spar/pro/RoleInTime); [folder "br"]: contains the data and provenance of the entities of type bibliographic resource (http://purl.org/spar/fabio/Expression); [folder "id"]: contains the data and provenance of the identifier entities (http://purl.org/spar/datacite/Identifier); [folder "ra"]: contains the data and provenance of the responsible agent type entities (http://xmlns.com/foaf/0.1/Agent); [folder "re"]: contains the data and provenance of resource embodiment entities (http://purl.org/spar/fabio/Manifestation) The inner folders are named through the supplier prefix of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to 06*0). After that, the folders have numeric names, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the zipped RDF data. At the same level, additional folders containing the provenance are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called prov, also in zipped JSON-LD format. For example, data related to the entity is located in the folder /br/06250/10000/1000/1000.zip, while information about provenance in /br/06250/10000/1000/prov/1000.zip This version of the dataset contains: 147,369,096 bibliographic entities 422,366,470 authors, 3,673,399 editors, and 115,504,540 publishers (counted by their roles, without disambiguating individual entities) 2,302,011 publication venues The compressed archives total ~43G, using the 7-zip compression algorithm, and expand to ~72G when decompressed on an ext4 filesystem. The JSON-LD files inside the archives are further compressed using the zip algorithm. It is recommended to process these inner files as compressed without extracting them, to manage data more efficiently. Additional information about OpenCitations Meta at the official webpage: https://opencitations.net/meta

提供机构:
Zenodo
创建时间:
2026-06-27
二维码
社区交流群
二维码
科研交流群
商业服务