M-DAIGT
收藏资源简介:
M-DAIGT是由多国研究机构联合构建的大规模AI生成文本检测基准数据集,涵盖新闻与学术两大关键领域。该数据集包含三万条平衡样本,其中人类撰写与AI生成文本各半,覆盖GPT-4、Claude等前沿模型生成的多样化内容。数据通过提取新闻标题与论文摘要作为提示词,采用角色随机分配策略生成风格各异的文本。该资源专为提升数字信息生态完整性而设计,支持跨领域AI文本检测模型的开发与评估,对防范虚假新闻与学术不端具有重要价值。
M-DAIGT is a large-scale benchmark dataset for AI-generated text detection, jointly constructed by research institutions across multiple countries. It covers two core domains: news and academia, and contains 30,000 balanced samples with an equal number of human-written and AI-generated texts. The dataset encompasses diverse content generated by cutting-edge models such as GPT-4 and Claude. Specifically, the texts are generated by using news headlines and paper abstracts as prompts, and adopting a random role assignment strategy to produce texts with distinct styles. This resource is specifically designed to enhance the integrity of the digital information ecosystem, supports the development and evaluation of cross-domain AI text detection models, and holds substantial value for preventing fake news and academic misconduct.
AI & Human Generated Text 数据集概述
数据集基本信息
- 许可证: MIT
- 语言: 英语
- 规模: 10K-100K
- 任务类别: 文本分类
数据集描述
- 全称: Artificial Intelligence Generated Abstracts (AI-GA)
- 内容构成: 包含摘要和标题
- 数据分布: 50%为AI生成摘要,50%为原始摘要
- 样本数量: 28,662个样本
- 样本结构: 每个样本包含摘要、标题和标签
标签说明
- 标签0: 原始摘要
- 标签1: AI生成摘要
技术特征
- AI生成摘要采用最先进的语言生成技术
- 主要基于GPT-3模型生成
应用用途
- 主要用于自然语言处理研究
- 特别适用于语言生成和机器学习实验
- 可用于AI文本检测研究
相关资源
- 原始GitHub仓库: https://github.com/panagiotisanagnostou/AI-GA
- 大型替代数据集: https://github.com/sakibsh/LLM

- 1M-DAIGT: A Shared Task on Multi-Domain Detection of AI-Generated Text卢森堡大学, 法赫德国王石油与矿产大学, 穆罕默德六世理工大学, 西迪·穆罕默德·本·阿卜杜拉大学 · 2025年



