遇见数据集

elisabeth-pl-pl/GRADTEX

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

GRADTEX 是一个用于细粒度AI生成文本检测的基准数据集。与传统的二元人类vs机器数据集不同,该数据集代表了机器参与文本生产的连续体:从纯人类撰写的文本,到经过语言模型编辑或部分重写的文本,再到完全从零生成的机器生成文本。数据集基于MAGE(Li等人,2024)构建,并在两个方向上进行了扩展:1)使用MAGE中过滤后的人类撰写文本作为所有编辑和部分生成场景的源材料;2)使用七种现代语言模型对这些文本应用十三种生成场景,产生不同机器参与程度的变体(例如,词级编辑、句级编辑、转述、补全、风格重写、基于主题的完全生成等)。其中九种场景用于训练和验证分割,其余四种仅出现在测试分割中(主要集中在Test C子集)。对于Test A(纯人类撰写文本 vs 纯机器生成文本),数据集还包含来自MAGE的3332个完全机器生成的文本,这些文本仅限于“topical_prompt”和“specified_prompt”生成设置。数据集旨在支持对机器参与文本的连续检测研究。

A benchmark for fine-grained AI-generated text detection. Unlike binary human-vs-machine datasets, this dataset represents the continuum of machine involvement in text production: from purely human-authored texts, through texts edited or partially rewritten by language models, to texts fully generated from scratch. The dataset is built on top of MAGE (Li et al., 2024) and extends it in two directions: 1) Filtered human-authored texts from MAGE are used as the source material for all editing and partial-generation scenarios. 2) Thirteen generation scenarios are applied to those texts using seven modern language models, producing variants with different degrees of machine involvement (token-level edits, sentence-level edits, paraphrases, completions, style rewrites, full topic-based generation). Nine of the thirteen scenarios are used in the train and validation splits; the remaining four appear only in the test split (concentrated in the Test C subset). For Test A (pure HWT vs pure MGT) the dataset also includes 3,332 fully machine-generated texts directly from MAGE, restricted to the `topical_prompt` and `specified_prompt` generation setups.

提供机构:
elisabeth-pl-pl
二维码
社区交流群
二维码
科研交流群
商业服务