相关数据集
MintVid
MintVid是由中国科学院与蚂蚁集团联合构建的高质量AI生成视频检测数据集,包含3000条视频样本,涵盖9种前沿生成模型。数据集分为三部分:1.5K高度逼真的专有模型生成视频(含文本/图像到视频内容)、2K基于3种公开模型的深度伪造视频,以及从短视频平台收集的真实场景含事实错误子集。其数据来源多样,覆盖通用内容、面部伪造和事实推理三大场景,旨在解决现有数据集时效性不足、时空一致性差等问题,为AI
arXiv2026-02-10 更新1260
"ChatGPT vs. Student: A Dataset for Source Classification of Computer Science Answers
A dataset comprising 500 data points was gathered by collecting answers to 250 computer science problems assigned in classes and quizzes from students. To generate this dataset, a response was s
DataCite Commons2023-04-12 更新110
Anonymous-2024/veracity-mfc-bench
--- dataset_info: features: - name: image_path dtype: string - name: RELEVANCY dtype: int64 - name: TOPIC dtype: int64 - name: DOCUMENT# dtype: string - name: evidence_id
Hugging Face2024-06-13 更新100
Generated and Real Academic Corpus for Evaluation (GRACE)
GRACE数据集是一个用于检测学术论文是由AI生成还是人类撰写的多语言数据集,包含英语和阿拉伯语的人类撰写和AI生成的论文。数据集的设计旨在确保内容的多样性和真实性,涵盖了不同的学术水平和文化背景。人类撰写的论文主要来源于语言评估考试如IELTS和TOEFL,而AI生成的论文则使用了多种先进的LLM模型生成。该数据集的应用领域主要集中在学术诚信和AI生成文本检测,旨在解决AI生成文本在学术环境中的
arXiv2024-12-24 更新560
AI text recognition dataset
We built an AI text recognition dataset that contains both human text and ChatGPT generated text in a total quantity of 35K. It contains three categories, namely news text, comment text, and question
科学数据银行2023-09-18 更新110



