Notice-based versus content-based detection of computer-generated content: cross-walk code, results, and figures
收藏资源简介:
Reproducibility package for an inter-method agreement study that cross-walks two independent paradigms for identifying computer-generated content (CGC) in the scholarly literature: a notice-based set (papers retracted with a Retraction Watch notice coded "Computer-Aided Content or Computer-Generated Content", 9,103 DOIs) and a content-based set (the union of the Problematic Paper Screener Tortured, SCIgen, and Suspect detectors, 30,600 DOIs). This deposit contains the cross-walk code (crosswalk.py, crosswalk_deep.py), the machine-readable results (crosswalk-results.json — every count in the associated manuscript reproduces from this file), and the four figures. The raw PPS detector exports and the Retraction Watch snapshot are third-party data and are not redistributed; their provenance, sizes, SHA-256 hashes, and acquisition dates are documented in README.md so the pipeline can be re-run against equivalent exports obtained from the original sources. Key result. The two methods agree on only 1,676 DOIs (18.4% of the notice set, 5.5% of the content set). The disagreement is structured: detectors differ sharply in precision/recall and cannot be pooled; content screening leads the retraction record by 20,622 not-yet-retracted papers. The backlog's headline 264,115 point-in-time citations is detector-skewed and false-positive-contaminated in the high-citation tail (the top item, 4,644 citations, is a legitimate Physiological Reviews review flagged by the Suspect heuristic only) and must be treated as a candidate signal, not a clean contamination count. This package supports a manuscript in preparation, proposed as a joint contribution with the PPS team (target: Quantitative Science Studies). It credits the PPS (Cabanac, Labbé, Magazinov) and the Retraction Watch Database (The Center for Scientific Integrity, ISSN 2692-4579) as data sources.



