C-ReD
收藏资源简介:
C-ReD是由清华大学等机构联合构建的中文AI生成文本检测基准数据集,涵盖新闻、问答、影评、作文及学术写作五大真实场景领域。数据集包含12,997条人工撰写文本和115,613条由9种大模型生成的AI文本,总规模达128,610条,数据来源于THUC-News、知乎、豆瓣等权威平台。通过精心设计的真实场景提示模板生成多领域文本,并经过自动化过滤与专家人工筛查双重质量控制。该数据集旨在解决中文AI文本检测中模型多样性不足、领域覆盖单一等核心问题,为检测算法提供跨领域、跨模型的评估基准。
C-ReD is a benchmark dataset for Chinese AI-generated text detection jointly constructed by Tsinghua University and other institutions. It covers five real-world scenario domains including news, question answering, movie reviews, student essays and academic writing. The dataset contains 12,997 manually authored texts and 115,613 AI-generated texts from 9 large language models, with a total of 128,610 instances. The data is sourced from authoritative platforms such as THUC-News, Zhihu and Douban. Multi-domain texts are generated via meticulously designed real-world prompt templates, and the dataset undergoes dual quality control including automated filtering and expert manual screening. This dataset aims to address core issues such as insufficient model diversity and narrow domain coverage in Chinese AI text detection, providing a cross-domain and cross-model evaluation benchmark for detection algorithms.
C-ReD 数据集概述
数据集状态
- 当前状态:即将发布。
数据集描述
- 描述:根据 README 文件内容,该数据集目前尚无具体信息提供,仅标注为“即将发布”。




