Data and code for Papers "AI use in manuscript preparation may widen the open-access citation advantage"
收藏资源简介:
This record contains the public data and code package for the study “AI use in manuscript preparation may widen the open-access citation advantage.” The archive includes redistributable paper-level metadata, AI-use disclosure labels, matched-control files, reference open-access citation metrics, cleaned intermediate outputs, supplementary tables, validation materials, and Python scripts used for the analysis. The workflow covers metadata assembly, broad keyword screening for AI-use disclosures, DeepSeek V4 Flash review, identity checks, reference open-access metric calculation, matched-control construction, cited journal-year control analysis, merged analysis, and statistical testing. The repository is organized into configuration files, numbered analysis scripts, archived reference outputs, workflow documentation, supplementary tables and extended methods, a file manifest with SHA256 checksums, and a package summary. The numbered scripts correspond to the numbered reference-output folders so that each analysis step can be traced to its archived output. The archive distributes public metadata, paper identifiers, article-level metadata fields, screening text snippets, AI-use disclosure evidence excerpts, reference open-access metrics, processed analysis results, and statistical outputs. Publisher full-text files, PDFs, XML files, HTML files, and other original full-text documents are not redistributed and remain on their original publication platforms. Required Python packages are listed in requirements.txt. The LLM review protocol used DeepSeek V4 Flash; no API credentials are included in this archive.



