遇见数据集

mohamedmady/Academic-Text-arxiv-gpt-gemini

收藏
Hugging Face2026-05-03 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个大规模学术领域基准,专门用于评估AI生成文本检测器在学术(STEM)领域实际分布偏移下的鲁棒性。数据集包含:469,008个人类撰写的arXiv段落(收集自2022年前发表的论文,确保纯人类生成文本),100,000个由GPT-3.5 Turbo生成的AI摘要,以及100,000个由Gemini 2.0 Flash生成的AI摘要,总计669,008个样本。AI文本通过官方API生成,主题与人类arXiv摘要分布仔细匹配,以确保语义和主题对齐紧密。

A large-scale academic-domain benchmark for AI-generated text detection. The dataset was specifically constructed to evaluate robustness of AI-text detectors under realistic distribution shift in the academic (STEM) domain. It contains: 469,008 human-written arXiv paragraphs (collected from papers published before 2022, guaranteeing purely human-generated text), 100,000 AI-generated abstracts produced by GPT-3.5 Turbo, and 100,000 AI-generated abstracts produced by Gemini 2.0 Flash, totaling 669,008 samples. The AI texts were generated using the official APIs on STEM topics carefully matched to the distribution of the human arXiv abstracts, ensuring close semantic and topical alignment.

提供机构:
mohamedmady
二维码
社区交流群
二维码
科研交流群
商业服务