遇见数据集

Anachronistic False Positives: A 26-Text Pre-LLM Canonical Benchmark for Commercial AI-Detection Tools

收藏
Zenodo2026-05-09 更新2026-05-26 收录
官方服务:

资源简介:

An open, reproducible benchmark for measuring false-positive behavior of commercial AI-detection tools on canonical human texts written before the existence of large language models. The corpus spans 81 BCE through 1962 CE and includes works in English, Middle English, Modern Chinese, and Classical Chinese. CORPUS 26 verbatim excerpts from canonical pre-LLM works, organized in three layers. Layer 1 — high-canonicity anchors (8 texts): Frost "The Road Not Taken" (1916); Thoreau Walden opening (1854); Hamlet "to be or not to be" soliloquy (~1600); Blake "The Tyger" (1794); 王勃 滕王阁序 (675 CE); 诸葛亮 出师表 (227 CE); 朱自清 荷塘月色 (1927); 范仲淹 岳阳楼记(1046). Layer 2 — same-era less-anthologized works (10 texts): Lincoln Cooper Union Address (1860); Eisenhower Farewell Address (1961); Wordsworth Tintern Abbey (1798); Henry V "Once more unto the breach" (~1599); 曹丕 典论·论文 (~220); 柳宗元 捕蛇者说 (~810); 鲁迅 阿长与山海经(1926); 王安石 游褒禅山记 (1054); Tennyson In Memoriam Canto VII (1850); 欧阳修 祭石曼卿文 (1067). Layer 3 — out-of-distribution probes (8 texts): Chaucer General Prologue in original Middle English (1387); Addison Spectator No. 1(1711); Macbeth "Tomorrow" soliloquy (1606); Dickens Tale of Two Cities opening (1859); Rossetti "Remember" (1862); 桓宽 盐铁论·本议 (81 BCE); 葛洪 抱朴子·内篇·畅玄 (~317); 王国维 人间词话 (1908). Each entry carries content-motif tags (death, war, oppression, grief, etc.) for downstream stratified analysis. BATTERY Five detectors run on every text: Sapling, Winston AI, GPTZero (commercial); HuggingFace RoBERTa-OpenAI-2019 (academic baseline via HF Inference Router); Claude Haiku 4.5 used as an LLM-as-judge proxy for the unofficial but widespread practice of educators pasting student work into a general-purpose chatbot. 130 detection calls in total. Raw per-call results, including refusal-state classification (hard content-filter, stop_reason=refusal, soft natural-language decline), are recorded in detection_results_classics.csv) A parallel smoke baseline of 8 modern-era texts (human academic writing 2008–2013 plus LLM-generated samples) is included in smoke_baseline/ so that the era-dependent inversion finding — particularly GPTZero classifying 100% of modern human writing as AI but 0% of pre-LLM canonical writing as AI — can be cross-checked against control data. HEADLINE FINDING One commercial detector (Sapling) classifies 24 of 26 pre-LLM canonical texts (92%) as AI-generated at the standard 0.5 threshold, including works predating large language models by intervals ranging from approximately 60 years (Frost, 1916) to over 1,300 years (王勃 滕王阁序, 675 CE) and over 2,100 years (桓宽 盐铁论, 81 BCE). GPTZero shows the opposite failure: zero false positives on canonical pre-LLM text but 100% false positives on modern human academic writing in the smoke baseline. The LLM-as-judge proxy refuses zero of 26 canonical inputs — including Hamlet's suicide soliloquy, Macbeth's nihilistic soliloquy, Cooper Union's slavery argumentation, and three elegiac texts (Tennyson, 欧阳修 祭石曼卿文, Rossetti) — but exhibits range-collapsed scoring (0.010–0.150 across canonical vs 0.72–0.92 across modern controls) that renders its numeric output classificatory rather than scalar. USE This dataset is intended as a public reference benchmark for evaluating any new AI-detection tool. Any tool that classifies a substantial fraction of these 26 canonical inputs as AI-generated must, by construction, be classifying something other than "machine origin". Researchers, institutional reviewers, and policy bodies are invited to reproduce the battery and add detectors and texts. FILES - historical_classics.csv — schema-aligned corpus (sample_id, category, source_model, text, length, ttr, sentence_count, language, era, author, title, source_provenance, motif_tags, expected_refused). - historical_classics/manifest.csv plus 26 individual .txt files — per-text metadata and verbatim source. - detection_results_classics.csv — 130-row raw battery output, one row per (sample, detector) pair. - smoke_baseline/ — 8-row modern control corpus and its 40-row detection results, used for the era-dependent inversion comparison. - figures/fig_classics_heatmap.{png,pdf} — 26 × 5 detection-probability matrix with ERR / REFUSED overlays. - figures/fig_classics_fp_bars.{png,pdf} — per-detector false-positive-rate bar chart (the headline figure). - 02_run_detection_battery.py, build_historical_classics.py, 04_plot_classics_results.py — reproduction scripts. CHANGELOG v1.1 (2026-05-09): added smoke baseline (n=8 modern control samples and 40 detection results); added pre-rendered figures (heatmap and FP bars) in PNG plus PDF; added figure-rendering script; corrected DOI in bundle README; metadata version bump. v1.0 (2026-05-08): initial release with 26 canonical texts, 5 detectors, 130 raw results, full reproduction scripts. COMPANION This dataset is the empirical core of an open letter to UNESCO's Section for Higher Education proposing replacement of automated AI-detection regimes with process-based academic-integrity verification. The letter and its technical report are deposited under separate.

提供机构:
Zenodo
创建时间:
2026-05-09
二维码
社区交流群
二维码
科研交流群
商业服务