Hybrid Source-Screening Benchmark for Problematic Scholarly Sources
收藏资源简介:
This repository contains the public reproducibility package supporting thestudy “Cost-Aware Selective Escalation for Screening Problematic ScholarlySources: A Controlled Hybrid Rule-LLM Benchmark.” The package includes a minimized benchmark of 1,031 scholarly records, savedpredictions from rule-based, LLM-only, selective-hybrid, non-gated hybrid, andtrainable machine-learning methods, canonical result tables, provenancemanifests, plotting data, and verification scripts. The benchmark contains 94 known retractions, 138 historical-watchlistassociations, and 226 composite should block cases. Raw third-party databaseexports, source abstracts, credentials, and free-form LLM reasoning are notincluded. The historical-watchlist label represents documented associationwith the historical source used in the study and should not be interpreted asa universal or current judgment of journal quality.



