Source Data for Manuscript: Identifying genomic data use with the Data Citation Explorer
收藏资源简介:
This page contains the source data for the manuscript describing the Data Citation Explorer, currently in review for publication. The preprint version can be found on this page. Files: DCE_manual_eval_sample.xlsx: This file was used to manually evaluate hits generated by the Data Citation Explorer. There are two separate sheets: one with publications returned by searches in PubMed and PubMed Central and another with publications returned by searches in Dimensions. Column descriptions can be found in the file itself. Each row in each evaluation sheet refers to a pair between a JAMO record and a linked publication. DCE_citation_report.tsv Contains JAMO record IDs and PubMed IDs from the initial 2020 DCE trial run. There are 238,994 unique JAMO IDs and 25,007 unique PubMed IDs. 76,511 JAMO records are linked with publications. Columns: jamo_id - unique JAMO record ID citation_count - Number of citations associated with each record citations - comma-delimited PubMed IDs for linked publications DCE_source_files.zip: This folder contains 3 files for each JAMO record in DCE_citation_report.tsv. For each JAMO record listed in the citation report, three files are provided: JAMO_ID_source.yaml - The fields extracted from the JAMO record that were relevant to the citation search, including any previously known PMIDs (manually curated). JAMO_ID_expand.yaml - The source record augmented with additional metadata discovered in other resources, including the citations that were discovered based on querying PubMed Central for the values in those metadata fields. JAMO_ID_audit.json - The audit path as a directed acyclic graph, in JSON. Of the ~238k JAMO records submitted to DCE, 6,979 contained anomalous fields that caused the records to be rejected for processing. This list is provided as NOT_PROCESSED.txt. Any records that were not processed are represented as zero-length files in the archive.



