遇见数据集

Radiation Oncology Open Access Journals Fabrication Candidate Dataset

收藏
Zenodo2026-05-23 更新2026-05-26 收录
官方服务:

资源简介:

In May 2026, Topaz and colleagues published an audit in The Lancet documenting a 12-fold rise in potentially fabricated references across 2·5 million biomedical papers following the November 2022 public release of ChatGPT. The finding implied that large language models had ushered in a systemic crisis of citation integrity for biomedical publishing. This project — led by Robert C. Miller, MD, MBA, FRSA (Mayo Clinic Emeritus / Indiana University School of Medicine) and Tessa Wrightson, DO (Marian University College of Osteopathic Medicine) — asked whether that population-level effect transported uniformly to a single, well-defined subspecialty. Radiation oncology was a useful test case: it is methodologically engaged, has strong technical-citation norms, and seven of its journals are fully open access with full-text JATS XML deposited in PubMed Central, making large-scale automated audit feasible. We replicated the Topaz analytic design across those seven Tier 1 open-access journals: Advances in Radiation Oncology, Clinical and Translational Radiation Oncology, Radiation Oncology, Physics and Imaging in Radiation Oncology, Technical Innovations and Patient Support in Radiation Oncology, the Journal of Radiation Research, and Reports of Practical Oncology and Radiotherapy. The audited corpus comprised 10,200 articles and 303,938 journal-type references published between 2005 and 2026. A Python pipeline performed automated reference verification through a three-step cascade: DOI lookup at Crossref, PMID lookup at PubMed, and Crossref bibliographic search by title and author. Three iterative debugging passes were required to produce a trustworthy verifier — a cascade-order fix preventing false flags on references from non-Crossref publishers, a PMID float-conversion guard, and a fuzzy-matcher threshold calibration aligning with Topaz’s Category 2 definition. The corrected pipeline identified 2,180 strong fabrication candidates (0·72% of references) distributed across Topaz’s three categories. Pre/post-ChatGPT comparison in the pre-specified balanced four-year window (2019–2022 versus 2023–2026) showed 12·96% of pre-ChatGPT articles and 13·65% of post-ChatGPT articles contained at least one candidate (Fisher exact: OR 0·94, p=0·45). No journal showed a statistically significant increase individually. Logistic regression on continuous proxy year showed a gradual long-term decline (OR 0·98 per year, p<0·001) with no inflection at the ChatGPT boundary. The analysis was powered to detect a 1·2-fold increase at 80%; a Topaz-scale 12-fold effect was firmly excluded. The radiation oncology literature shows no LLM-era citation fabrication surge. This finding does not refute Topaz’s population-level result but specifies its non-uniform distribution across biomedicine. Subspecialty audits should precede field-wide intervention; an editorial response calibrated to the population mean may over- or under-serve fields that lie far from it. The work also demonstrates that automated reference-integrity audit is tractable for any subspecialty whose journals deposit JATS XML in PubMed Central, opening a path to routine surveillance rather than one-off detection. The project produced a submitted Correspondence titled “The post-ChatGPT citation fabrication surge is not universal across medical specialties”, a nine-section methodological web appendix, the 328,091-reference dataset, the Python pipeline source code, a standalone 8-page PDF methodology explainer, and a Zenodo deposit (DOI: 10.5281/zenodo.20350450) carrying both code and data for permanent citation. The code is also released on GitHub under MIT license.

提供机构:
Zenodo
创建时间:
2026-05-23
二维码
社区交流群
二维码
科研交流群
商业服务