遇见数据集

Anachronistic False Positives: A 26-Text Pre-LLM Canonical Benchmark for Commercial AI-Detection Tools

收藏
Zenodo2026-05-09 更新2026-05-26 收录
官方服务:

资源简介:

<p>An open, reproducible benchmark for measuring false-positive behavior of commercial AI-detection tools on canonical human texts written before the existence of large language models. The corpus spans 81 BCE through 1962 CE and includes works in English, Middle English, Modern Chinese, and Classical Chinese.</p> <p><strong>Corpus.</strong> 26 verbatim excerpts from canonical pre-LLM works, organized in three layers: (L1) high-canonicity anchors (Frost, Thoreau, Hamlet, Blake, 滕王阁序, 出师表, 荷塘月色, 岳阳楼记); (L2) same-era less-anthologized works (Cooper Union, Eisenhower Farewell, Wordsworth's Tintern Abbey, Henry V, 典论·论文, 捕蛇者说, 阿长与山海经, 王安石 游褒禅山记, Tennyson In Memoriam, 欧阳修 祭石曼卿文); (L3) out-of-distribution probes (Chaucer's Middle English, Addison's Spectator, Macbeth, Dickens, Rossetti, 盐铁论, 抱朴子, 王国维 人间词话). Each entry carries content-motif tags (death, war, oppression, grief, etc.) for downstream stratified analysis.</p> <p><strong>Battery.</strong> Five detectors run on every text: Sapling, Winston AI, GPTZero (commercial); HuggingFace RoBERTa-OpenAI-2019 (academic baseline via HF Inference Router); Claude Haiku 4.5 used as an LLM-as-judge proxy for the "shadow detection regime" that occurs when teachers paste student work into a general-purpose chatbot. 130 detection calls total. Raw per-call results, including refusal-state classification (hard content-filter, stop_reason=refusal, soft natural-language decline), are recorded in <code>detection_results_classics.csv</code>.</p> <p><strong>Headline finding.</strong> One commercial detector (Sapling) classifies 24 of 26 pre-LLM canonical texts (92%) as AI-generated at threshold 0.5, including works predating LLMs by intervals ranging from approximately 60 years (Frost, 1916) to over 1,300 years (王勃 滕王阁序, 675 CE) and over 2,100 years (桓宽 盐铁论, 81 BCE). GPTZero shows the opposite failure: zero false positives on canonical pre-LLM text but 100% false positives on modern human academic writing in the smoke set. The LLM-as-judge proxy refuses zero of 26 canonical inputs (including Hamlet's suicide soliloquy, Macbeth's nihilism, Cooper Union's slavery argumentation, and three elegiac texts) but exhibits range-collapsed scoring (0.010-0.150) that renders its numeric output classificatory rather than scalar.</p> <p><strong>Use.</strong> This dataset is intended as a public reference benchmark for evaluating any new AI-detection tool. Any tool that classifies a substantial fraction of these 26 canonical inputs as AI-generated must, by construction, be classifying something other than "machine origin". Researchers, institutional reviewers, and policy bodies are invited to reproduce the battery and add detectors / texts.</p> <p><strong>Files.</strong></p> <ul> <li><code>historical_classics.csv</code> — schema-aligned corpus (sample_id, category, source_model, text, length, ttr, sentence_count, language, era, author, title, source_provenance, motif_tags, expected_refused)</li> <li><code>historical_classics/manifest.csv</code> + 26 individual <code>.txt</code> files — per-text metadata and verbatim source</li> <li><code>detection_results_classics.csv</code> — 130-row raw battery output, one row per (sample, detector)</li> <li><code>02_run_detection_battery.py</code> + <code>build_historical_classics.py</code> — reproduction scripts</li> </ul> <p><strong>Companion.</strong> This dataset is the empirical core of an open letter to UNESCO's Section for Higher Education proposing replacement of automated AI-detection regimes with process-based academic-integrity verification. The letter and its technical report are deposited under separate Zenodo records.</p>

提供机构:
Zenodo
创建时间:
2026-05-09
二维码
社区交流群
二维码
科研交流群
商业服务