BanglaAiDetect: A Large-Scale Benchmark Dataset for Human vs AI-Generated Bangla Text
收藏资源简介:
OverviewBanglaAiDetect is a balanced dataset of 40,000 Bangla texts, evenly split between human-written and AI-generated content. It is intended as a resource for research on AI text detection, dataset analysis, and other natural language processing (NLP) tasks in Bangla. Data Composition Human-Written Texts: Collected from publicly available sources including Bengali newspapers, blogs, and social media posts. Personally identifiable information was removed where necessary. AI-Generated Texts: Gemini 2.5 Pro outputs taken from the BanSum dataset (Hasan et al., 2024) [CC BY 4.0, DOI: 10.17632/rxhj7g6y2k.1]. GPT-4 and Grok-4 outputs generated separately using more than 2,000 unique prompts, then filtered for length and quality. Key Features Total size: 40,000 samples (20,000 human / 20,000 AI). Balanced across classes. Includes outputs from three different LLMs (Gemini 2.5 Pro, GPT-4, Grok-4).



