AH&AITD – Arslan's Human and AI Text Database
收藏资源简介:
AH&AITD is a comprehensive benchmark dataset designed to support the evaluation of AI-generated text detection tools. The dataset contains 11,580 samples spanning both human-written and AI-generated content across multiple domains. It was developed to address limitations in previous datasets, particularly in terms of diversity, scale, and real-world applicability. To facilitate research in the detection of AI-generated text by providing a diverse, multi-domain dataset. This dataset enables fair benchmarking of detection tools across various writing styles and content categories. Composition 1. Human-Written Samples (Total: 5,790) Collected from: Open Web Text (2,343 samples) Blogs (196 samples) Web Text (397 samples) Q&A Platforms (670 samples) News Articles (430 samples) Opinion Statements (1,549 samples) Scientific Research Abstracts (205 samples) 2. AI-Generated Samples (Total: 5,790) Generated using: ChatGPT (1,130 samples) GPT-4 (744 samples) Paraphrase Models (1,694 samples) GPT-2 (328 samples) GPT-3 (296 samples) DaVinci (GPT-3.5 variant) (433 samples) GPT-3.5 (364 samples) OPT-IML (406 samples) Flan-T5 (395 samples)



