Zenodo Open Metadata Snapshot. Training Dataset For Records Classifier Building
收藏资源简介:
This dataset contains Zenodo's published open access record's metadata as of 6th of March 2017.<br> <br> It's composed of: A TAR.GZ archive <strong>zenodo_open_metadata_17_05_2018.tar.gz</strong> containing the full dataset: Full dataset contains: Metadata of 404648 Zenodo records. Metadata of 8472 records which were classified as SPAM records by Zenodo staff. Dataset contains only already publicly available metadata of all of the records. In two cases, the metadata has been altered: One title from a SPAM-labelled record has been altered as it contained an e-mail address. One SPAM-labelled record has been removed from the full dataset Data format description: Dataset is a JSON file, containing a single list of 176741 key-value dictionaries.<br> <br> Each dictionary contains the terms:<br> <strong>part_of, thesis, description, doi, meeting, imprint, references, recid, alternate_identifiers, resource_type, journal, related_identifiers, title, subjects, notes, creators, communities, access_right, keywords, contributors, publication_date</strong><br> <br> which are corresponding to the fields with the same name available in Zenodo's record jsonschema v1.0.0: https://github.com/zenodo/zenodo/blob/master/zenodo/modules/records/jsonschemas/records/record-v1.0.0.json In addition, some terms have been altered:<br> <br> The term <strong>files</strong> contains a list of dictionaries containing <strong>filetype</strong>, <strong>size</strong> and <strong>filename </strong>only.<br> The term <strong>license</strong> contains a short Zenodo ID of the license (e.g "cc-by").<br> The term <strong>spam</strong> contains a boolean value, determining whether given record was marked as SPAM record by Zenodo staff.<br> <br> Some values for the top-level terms, which were missing in the metadata may contain a <strong>null</strong> value.



