DeHAT v1: A German Multi-Register Corpus of Human and AI-Generated Text Variants
收藏资源简介:
DeHAT is a German corpus for comparing human-written texts with texts produced by language models. Each human text provides the starting point for a text family. Claude Opus 4.7 derived a writing prompt from it. Eleven models received that prompt and produced new texts without seeing the original. GPT-5.5, one of these eleven models, also produced a version by directly revising the original. This structure supports comparisons of word choice, sentence patterns and formatting across models and text categories. The public edition provides texts and writing prompts from 16,519 text families, 18 sources and nine text categories. The 226,092 texts and writing prompts cover news, academic writing, law, parliamentary debate and other kinds of German writing. Explore and use the corpus. Start with dehat-v1.0.0.zip. Its local browser viewer lets you read texts side by side and filter by source, text category or dataset split (training, development or test). The viewer requires Python 3.10 or later. Data are supplied in JSONL and Parquet, with an export tool for TXT, CSV and CWB VRT. Download and extract the complete archives, as Zenodo cannot preview their contents. The project website provides instructions in German and English. Access and reuse. Source rights determine which members of each text family are included as full text, available for local reconstruction or documented only through metadata. SOURCE_RIGHTS.md and the accompanying metadata specify rights and attribution for human texts and model outputs. The reconstruction tools add further human texts and Europarl direct revisions locally and check that they match the recorded versions. The README gives the full breakdown of available texts, metadata and reconstruction options. Analyses and code. dehat-v1.0.0-code.zip contains executable analyses, 30 result tables and documentation of how each model was prompted, how its responses were checked and when a new response was requested. These analyses use 16,310 complete text families selected for the study. Repeating the analyses requires an authorized copy of that set, verified using the supplied checksums. A separate supplement contains scores for 1,300 texts evaluated with the Grammarly AI detector and a standalone analysis script. It includes neutral identifiers for comparing scores without distributing the evaluated texts or screenshots. The public corpus can be used independently for new analyses.



