Data and code for Papers disclosing AI use cite more open-access literature
收藏资源简介:
This record contains the public data and code package for the study “Papers disclosing AI use cite more open-access literature”. The archive includes redistributable paper-level metadata, AI-use disclosure labels, matched-control files, reference open-access citation metrics, cleaned intermediate outputs, statistical result tables, validation materials, and Python scripts used for the analysis. The workflow covers metadata assembly, broad keyword screening for AI-use disclosures, DeepSeek V4 Flash review, identity checks, reference open-access metric calculation, matched-control construction, merged analysis, and statistical testing. The repository is organized into configuration files, analysis scripts, reference outputs, workflow documentation, a manifest with file checksums, and a package summary. The numbered scripts correspond to the numbered reference-output folders so that each analysis step can be traced to its archived output. The archive distributes public metadata, paper identifiers, article-level metadata fields, screening text snippets, AI-use disclosure evidence excerpts, reference open-access metrics, processed analysis results, and statistical outputs. Publisher full-text files, PDFs, XML files, HTML files, and other original full-text documents are not redistributed and remain on their original publication platforms. Required Python packages are listed in requirements.txt. The LLM review step uses DeepSeek V4 Flash; users rerunning that step should provide their own API credentials.



