Citation Integrity in the LLM Era: A Complete-Corpus Audit of Clinical and Translational Radiation Oncology
收藏资源简介:
In May 2026 a Lancet audit by Topaz and colleagues reported a 12-fold rise in potentially fabricated references across roughly 2·5 million biomedical papers following the November 2022 public release of ChatGPT. The result implied that large language models had introduced a systemic citation-integrity problem for biomedical publishing and prompted calls for field-wide editorial intervention. Whether that population-level finding transports to any individual subspecialty journal remained an empirical question. This project — led by Jennifer Marie Ritchie, BA, and Robert C. Miller, MD, MBA, FRSA (Indiana University School of Medicine and Mayo Clinic Emeritus Office) — examined whether the post-ChatGPT citation-integrity signal could be detected in a single, well-defined open-access radiation oncology journal. Clinical and Translational Radiation Oncology was a natural test case: an ESTRO-portfolio Elsevier journal launched in December 2016 with full reference deposition in PubMed Central, complete first-decade coverage, and an editorial scope that spans the methodologically engaged translational interface where citation precision matters most. Every article indexed in PubMed Central under the journal's MEDLINE abbreviation through the May 2026 data freeze was enumerated through the NCBI E-utilities API and processed end-to-end through a custom Python verification pipeline. JATS-formatted full-text XML was parsed to extract every reference and its identifiers; each reference was then routed through a three-step verification cascade — DOI lookup at Crossref, PubMed identifier lookup at the National Library of Medicine, and a fallback bibliographic title-and-author search at Crossref. Title-match similarity was computed with RapidFuzz token-set scoring; matches required concordance on title, first-author surname, and publication year. References were stratified by publication era using the maximum cited reference year per article as a strict lower bound proxy for submission date. Pre-LLM, transitional, and post-LLM cohorts were defined to align with the November 2022 ChatGPT boundary used by Topaz, with 2022-boundary articles handled as a separate transition group. The strictest fabrication-prior subset — references for which automated lookup returned no matching record after exhaustive title-and-author search — served as the primary outcome. Era differences were tested with Fisher’s exact at the reference level and a cluster-aware Mann–Whitney U test at the per-article level, with a pre-specified outlier-sensitivity analysis to address the dominant influence that any single systematic review can exert on a reference-level proportion test. References that the automated cascade could not confirm — those with no matching record found and those with insufficient bibliographic information — were taken forward for manual dual-reviewer adjudication as a complete census rather than a sample. Two reviewers independent of the authorship team classified every unconfirmable reference into four pre-specified categories: legitimate non-indexed sources, legitimate journal articles carrying citation errors, untraceable but plausibly legitimate, or fabricated. Cohen's κ and raw percent agreement quantified inter-rater reliability; disagreements were resolved by discussion or by a third adjudicator. The project produced a submitted Short Communication for Clinical and Translational Radiation Oncology, a reference-level dataset, the Python verification pipeline released under MIT licence, three publication-quality figures, a manual-adjudication scoring workbook, ICMJE disclosures for both authors, a standalone methodology PDF, and a Zenodo deposit carrying code and data for permanent citation. The pipeline is portable to any radiation oncology or biomedical journal that deposits full-text XML in PubMed Central.



