Supporting data for Measuring the State of Open Science in Transportation Using Large Language Models
收藏资源简介:
This record contains the supporting data for the paper "Measuring the State of Open Science in Transportation Using Large Language Models." The study uses large language models (LLMs) to extract open science indicators, such as data use, data availability, and code availability, from the full text of 10,724 research articles published from 2019 to 2024 in the seven Transportation Research journals (Parts A through F and Interdisciplinary Perspectives). The release includes three datasets, which together allow others to test the extraction pipeline, reproduce the analysis, and assess the accuracy of the LLM outputs. 1. Sample full-text data (sample_xml_data.zip)Five full-text articles from Transportation Research Part A: Policy and Practice in the XML format used as input to the extraction pipeline. Each file is named by the article's DOI. All five are open-access articles published under the Creative Commons Attribution 4.0 license (CC BY 4.0), which permits redistribution. The sample shows the expected input format and lets users run the pipeline end to end. 2. Analysis dataset (analysis_dataset.zip)The LLM-extracted features and derived variables used in the paper's analysis.- papers_cleaned.csv: one row per article (10,724 articles, 76 columns). It includes bibliographic metadata (journal, year, DOI, number of authors, primary institution, region, citation count, review time), LLM-extracted indicators (e.g., whether data or code is used, whether data is cited or linked, whether a data repository or public code is available, availability statements, and software and hardware reported).- papers_cleaned_with_new_columns.csv: the same table with an added indicator for whether the study is quantitative.- papers_cleaned_p2.csv: the same table with added code link checks: whether the code link is live, whether the code is hosted on GitHub, and whether the GitHub repository has a README.- datasets_aggregated.csv: one row per unique dataset URL cited. It records the hosting domain, dataset type, source category, number of papers using it, years of first and last use, and associated regions.- paper_dataset_mapping.csv: links each article DOI to the dataset URLs it references. 3. Manual validation dataset (manual_validation_dataset.zip)Human labels used to evaluate the LLM extraction. Two annotators independently labeled articles for eight indicators: quantitative study, data used, data cited, data repository available, data link valid, code publicly available, code link valid, and simulation study. h1_96.csv contains the first annotator's labels for 113 articles, and h2_96.csv contains the second annotator's labels for 96 of those articles. Together they support both LLM-versus-human and inter-annotator agreement analysis. The articles span all seven journals. The code for the extraction pipeline and analysis is available at https://github.com/rrinTransportation/OpenMOST. License: The analysis and validation datasets are released under CC BY 4.0. The sample articles remain under their original CC BY 4.0 licenses.



