Citation Integrity in the LLM Era: A Complete-Corpus Audit of Clinical and Translational Radiation Oncology
收藏资源简介:
In May 2026 a Lancet audit by Topaz and colleagues reported a 12-fold rise in potentially fabricated references across roughly 2·5 million biomedical papers following the November 2022 public release of ChatGPT. The result implied that large language models had introduced a systemic citation-integrity problem for biomedical publishing and prompted calls for field-wide editorial intervention. Whether that population-level finding transports to any individual subspecialty journal remained an empirical question. This project — led by Jennifer Marie Ritchie, BA, and Robert C. Miller, MD, MBA, FRSA (Indiana University School of Medicine and Mayo Clinic Emeritus Office) — examined whether the post-ChatGPT citation-integrity signal could be detected in a single, well-defined open-access radiation oncology journal. Clinical and Translational Radiation Oncology was a natural test case: an ESTRO-portfolio Elsevier journal launched in December 2016 with full reference deposition in PubMed Central, complete first-decade coverage, and an editorial scope that spans the methodologically engaged translational interface where citation precision matters most. Every article indexed in PubMed Central under the journal's MEDLINE abbreviation through the May 2026 data freeze was enumerated through the NCBI E-utilities API and processed end-to-end through a custom Python verification pipeline. JATS-formatted full-text XML was parsed to extract every reference and its identifiers; each reference was then routed through a three-step verification cascade — DOI lookup at Crossref, PubMed identifier lookup at the National Library of Medicine, and a fallback bibliographic title-and-author search at Crossref. Title-match similarity was computed with RapidFuzz token-set scoring; matches required concordance on title, first-author surname, and publication year. References were stratified by publication era using the maximum cited reference year per article as a strict lower bound proxy for submission date. Pre-LLM, transitional, and post-LLM cohorts were defined to align with the November 2022 ChatGPT boundary used by Topaz, with 2022-boundary articles handled as a separate transition group. The strictest fabrication-prior subset — references for which automated lookup returned no matching record after exhaustive title-and-author search — served as the primary outcome. Era differences were tested with Fisher’s exact at the reference level and a cluster-aware Mann–Whitney U test at the per-article level, with a pre-specified outlier-sensitivity analysis to address the dominant influence that any single systematic review can exert on a reference-level proportion test. References that the automated cascade could not confirm — those with no matching record found and those with insufficient bibliographic information — were taken forward for manual dual-reviewer adjudication as a complete census rather than a sample. Two reviewers independent of the authorship team classified every unconfirmable reference into four pre-specified categories: legitimate non-indexed sources, legitimate journal articles carrying citation errors, untraceable but plausibly legitimate, or fabricated. Cohen's κ and raw percent agreement quantified inter-rater reliability; disagreements were resolved by discussion or by a third adjudicator. The project produced a submitted Short Communication for Clinical and Translational Radiation Oncology, a reference-level dataset, the Python verification pipeline released under MIT licence, three publication-quality figures, a manual-adjudication scoring workbook, ICMJE disclosures for both authors, a standalone methodology PDF, and a Zenodo deposit carrying code and data for permanent citation. The pipeline is portable to any radiation oncology or biomedical journal that deposits full-text XML in PubMed Central.
2026年5月,托帕兹(Topaz)及其同事在《柳叶刀》(Lancet)发表的一项审计研究显示,自2022年11月ChatGPT公开发布以来,约250万篇生物医学论文中潜在伪造参考文献的数量增长了12倍。该结果表明,大语言模型(Large Language Model, LLM)给生物医学出版带来了系统性的引用完整性问题,也引发了学界开展全领域编辑干预的呼吁。而这一群体层面的发现是否适用于任意单个专科期刊,仍是一个有待实证解答的问题。 本项目由文学学士珍妮弗·玛丽·里奇(Jennifer Marie Ritchie)、医学博士/工商管理硕士/皇家艺术学会会士罗伯特·C·米勒(Robert C. Miller,印第安纳大学医学院及梅奥诊所荣休办公室)领衔,旨在探究ChatGPT发布后的引用完整性异常信号,能否在单一定义明确的开放获取放射肿瘤学期刊中被检测到。《临床与转化放射肿瘤学(Clinical and Translational Radiation Oncology)》是理想的测试案例:它是欧洲放射治疗与肿瘤学会(European Society for Radiotherapy and Oncology, ESTRO)旗下爱思唯尔(Elsevier)期刊,于2016年12月创刊,所有参考文献均提交至PubMed Central(PubMed中央库),拥有完整的首十年出版数据覆盖范围,其编辑范围涵盖方法学严谨的转化研究前沿——这一领域对引用精准度要求极高。 研究通过NCBI E-utilities API枚举了截至2026年5月数据冻结期内,以该期刊MEDLINE缩写在PubMed Central中收录的所有文章,并通过自研Python验证流水线进行全流程处理。首先解析JATS格式的全文XML以提取所有参考文献及其标识符;随后将每篇参考文献纳入三级验证级联流程:首先在Crossref平台进行DOI检索,其次在美国国家医学图书馆进行PubMed标识符检索,若前两步失败,则在Crossref平台通过参考文献标题与作者信息进行兜底检索。使用RapidFuzz的Token-set评分计算标题匹配相似度,匹配需同时满足标题、第一作者姓氏及出版年份一致。 以每篇文章被引参考文献的最大出版年份作为投稿日期的严格下界代理变量,将参考文献按出版时代分层。参考托帕兹团队采用的2022年11月ChatGPT发布时间节点,将研究时段划分为大语言模型前、过渡期及大语言模型后三个队列,2022年边界所在的文章单独作为过渡组。本研究的主要结局指标为最严格的伪造前置子集:即经过穷尽式标题与作者检索后,自动检索未返回匹配记录的参考文献。不同时代的差异分别在参考文献层面采用费希尔精确检验(Fisher’s exact test),在单篇文章层面采用聚类校正的曼-惠特尼U检验(cluster-aware Mann–Whitney U test),并预先设定了异常值敏感性分析,以应对任何单篇系统综述对参考文献层面比例检验造成的主导性影响。 所有自动验证流程无法确认的参考文献(包括未检索到匹配记录的条目,以及书目信息不足的条目)将作为完整普查而非抽样,交由两名独立于作者团队的审稿人进行双人手工裁定。两名审稿人将每篇无法确认的参考文献划分为四个预先设定的类别:合法未索引来源、存在引用错误的合法期刊文章、无法溯源但看似合理的来源,或伪造参考文献。采用科恩κ系数(Cohen's κ)及原始百分比一致性量化评阅者间信度,分歧通过讨论或第三名裁定者解决。 本项目已向《临床与转化放射肿瘤学》提交短篇通讯(Short Communication)文章,同时产出了参考文献层面的数据集、以MIT许可证开源的Python验证流水线、三张可用于出版的高质量图表、手工裁定评分工作簿、两位作者的国际医学期刊编辑委员会(International Committee of Medical Journal Editors, ICMJE)披露声明、独立的方法学PDF文档,以及一个可永久引用的Zenodo存档库,其中包含代码与数据。该流水线可移植至所有将全文XML提交至PubMed Central的放射肿瘤学或生物医学期刊。



