遇见数据集

Radiation Oncology Open Access Journals Fabrication Candidate Dataset

收藏
Zenodo2026-05-25 更新2026-05-26 收录
官方服务:

资源简介:

In May 2026, Topaz and colleagues published an audit in The Lancet documenting a 12-fold rise in potentially fabricated references across 2·5 million biomedical papers following the November 2022 public release of ChatGPT. The finding implied that large language models had ushered in a systemic crisis of citation integrity for biomedical publishing. This project — led by Robert C. Miller, MD, MBA, FRSA (Mayo Clinic Emeritus / Indiana University School of Medicine) and Tessa Wrightson, DO (Marian University College of Osteopathic Medicine) — asked whether that population-level effect transported uniformly to a single, well-defined subspecialty. Radiation oncology was a useful test case: it is methodologically engaged, has strong technical-citation norms, and seven of its journals are fully open access with full-text JATS XML deposited in PubMed Central, making large-scale automated audit feasible. We replicated the Topaz analytic design across those seven Tier 1 open-access journals: Advances in Radiation Oncology, Clinical and Translational Radiation Oncology, Radiation Oncology, Physics and Imaging in Radiation Oncology, Technical Innovations and Patient Support in Radiation Oncology, the Journal of Radiation Research, and Reports of Practical Oncology and Radiotherapy. The audited corpus comprised 10,200 articles and 303,938 journal-type references published between 2005 and 2026. A Python pipeline performed automated reference verification through a three-step cascade: DOI lookup at Crossref, PMID lookup at PubMed, and Crossref bibliographic search by title and author. Three iterative debugging passes were required to produce a trustworthy verifier — a cascade-order fix preventing false flags on references from non-Crossref publishers, a PMID float-conversion guard, and a fuzzy-matcher threshold calibration aligning with Topaz’s Category 2 definition. The project produced a submitted Correspondence titled “The post-ChatGPT citation fabrication surge is not universal across medical specialties”, a nine-section methodological web appendix, the 328,091-reference dataset, the Python pipeline source code, a standalone 8-page PDF methodology explainer, and a Zenodo deposit (DOI: 10.5281/zenodo.20350450) carrying both code and data for permanent citation. The code is also released on GitHub under MIT license

提供机构:
Zenodo
创建时间:
2026-05-23
二维码
社区交流群
二维码
科研交流群
商业服务