Data Citation Corpus Data File
收藏资源简介:
Data file for the first release of the Data Citation Corpus, produced by DataCite and Make Data Count as part of an ongoing grant project funded by the Wellcome Trust. Read more about the project. The data file includes 10,006,058 data citation records in JSON and CSV formats. The JSON file is the version of record. The data citations in the file originate from DataCite Event Data and a project by Chan Zuckerberg Initiative (CZI) to identify mentions to datasets in the full text of articles. Each data citation record is comprised of: A pair of identifiers: An identifier for the dataset (a DOI or an accession number) and the DOI of the publication object (journal article or preprint) in which the dataset is cited Metadata for the cited dataset and for the citing publication object The data file includes the following fields: Field Description Required? id Internal identifier for the citation Yes created Date of item's incorporation into the corpus Yes updated Date of item's most recent update in corpus Yes repository Repository where cited data is stored No publisher Publisher for the article citing the data No journal Journal for the article citing the data No title Title of cited data No objId DOI of article where data is cited Yes subjId DOI or accession number of cited data Yes publishedDate Date when citing article was published No accessionNumber Accession number of cited data No doi DOI of cited data No relationTypeId Relation type in metadata between citation object and subject No source Source where citation was harvested Yes subjects Subject information for cited data No affiliations Affiliation information for creator of cited data No funders Funding information for cited data No Additional documentation about the citations and metadata in the file is available on the Make Data Count website. Feedback on the data file can be submitted via Github. For general questions, email info@makedatacount.org.
本数据集为数据引用语料库(Data Citation Corpus)首次发布的数据文件,由DataCite与Make Data Count联合打造,作为惠康信托基金会(Wellcome Trust)资助的一项持续性资助项目的组成部分。可查阅该项目的详细介绍。 本数据文件共计包含10,006,058条数据引用记录,支持JSON与CSV两种格式,其中JSON文件为正式发布版本。 本文件中的数据引用源自DataCite事件数据(DataCite Event Data),以及陈·扎克伯格倡议(Chan Zuckerberg Initiative, CZI)的一项项目——该项目旨在从学术论文全文中识别对数据集的提及内容。 单条数据引用记录由以下两部分构成: 1. 标识符对:分别为被引用数据集的标识符(DOI或登录号),以及引用该数据集的出版载体(期刊论文或预印本)的DOI; 2. 被引用数据集与引用出版载体的元数据。 本数据文件包含以下字段: | 字段名 | 字段描述 | 是否必填 | |-------|----------|----------| | id | 该数据引用的内部标识符 | 是 | | created | 该条目被纳入语料库的日期 | 是 | | updated | 该条目在语料库中最近一次更新的日期 | 是 | | repository | 存储被引用数据的仓储 | 否 | | publisher | 引用该数据的论文的出版方 | 否 | | journal | 引用该数据的论文所属期刊 | 否 | | title | 被引用数据的标题 | 否 | | objId | 引用数据的论文的DOI | 是 | | subjId | 被引用数据的DOI或登录号 | 是 | | publishedDate | 引用论文的发表日期 | 否 | | accessionNumber | 被引用数据的登录号 | 否 | | doi | 被引用数据的DOI | 否 | | relationTypeId | 元数据中引用对象与被引用对象之间的关系类型标识 | 否 | | source | 该数据引用的采集来源 | 是 | | subjects | 被引用数据的主题信息 | 否 | | affiliations | 被引用数据创作者的所属机构信息 | 否 | | funders | 被引用数据的资助方信息 | 否 | 有关本文件中数据引用与元数据的补充文档,可于Make Data Count官网获取。若您对本数据文件有反馈意见,可通过Github提交;一般性疑问请发送邮件至info@makedatacount.org。



