Generated and Real Academic Corpus for Evaluation (GRACE)
收藏资源简介:
GRACE数据集是一个用于检测学术论文是由AI生成还是人类撰写的多语言数据集,包含英语和阿拉伯语的人类撰写和AI生成的论文。数据集的设计旨在确保内容的多样性和真实性,涵盖了不同的学术水平和文化背景。人类撰写的论文主要来源于语言评估考试如IELTS和TOEFL,而AI生成的论文则使用了多种先进的LLM模型生成。该数据集的应用领域主要集中在学术诚信和AI生成文本检测,旨在解决AI生成文本在学术环境中的滥用问题。
The GRACE dataset is a multilingual dataset for detecting whether academic papers are AI-generated or human-written. It contains human-written and AI-generated papers in English and Arabic. The dataset is designed to ensure content diversity and authenticity, covering different academic levels and cultural backgrounds. The human-written papers are mainly sourced from language proficiency tests such as IELTS and TOEFL, while the AI-generated papers are created using multiple advanced large language models (LLMs). The primary application fields of this dataset focus on academic integrity and AI-generated text detection, aiming to address the abuse of AI-generated texts in academic environments.

- 1GenAI Content Detection Task 2: AI vs. Human -- Academic Essay Authenticity Challenge卡塔尔计算研究所, 卡塔尔大学, TOBB ETU, 哈马德·本·哈利法大学 · 2024年



