CHALLENGE-IT
收藏资源简介:
CHALLENGE-IT是一个专门用于检测大型语言模型(LLM)生成文本的意大利语基准数据集,旨在为评估和提升AI生成文本检测模型在意大利语上的性能提供一个具有挑战性的测试平台。数据集包含总计360,438个文本样本,来源于四个意大利语语料库(CHANGE-IT, C4-IT, PaISà),并由14个不同的大型语言模型生成。数据集被划分为训练集(298,472个样本)、验证集(52,672个样本)以及一个特别设计的测试集(9,294个样本)。这个测试集是一个‘硬负例’子集,其样本对于基于RoBERTa和EuroBERT的分类器都同时难以区分,旨在模拟现实世界中最具挑战性的检测场景。在该测试集上,所有已评估系统的AUC得分在0.39到0.56之间,接近随机猜测水平,突显了其难度和作为严格基准的价值。该数据集适用于文本分类任务,特别是AI生成文本检测领域的研究和模型开发。
CHALLENGE-IT is a specialized benchmark dataset for detecting text generated by large language models (LLMs) in Italian. It aims to provide a challenging testing platform for evaluating and improving AI-generated text detection models in Italian. The dataset contains a total of 360,438 text samples, sourced from four Italian corpora (CHANGE-IT, C4-IT, PaISà) and generated by 14 different large language models. It is divided into a training set (298,472 samples), a validation set (52,672 samples), and a specially designed test set (9,294 samples). This test set is a hard negative subset where samples are simultaneously difficult to distinguish for classifiers based on RoBERTa and EuroBERT, designed to simulate the most challenging detection scenarios in the real world. On this test set, all evaluated systems achieve AUC scores between 0.39 and 0.56, close to random guessing, highlighting its difficulty and value as a rigorous benchmark. The dataset is suitable for text classification tasks, particularly in the field of AI-generated text detection research and model development.
数据集概述:CHALLENGE-IT
- 名称:CHALLENGE-IT(Italian Benchmark for LLM-Generated Text Detection)
- 语言:意大利语
- 许可证:CC-BY-4.0
- 任务类别:文本分类
- 大小规模:100K < n < 1M
- 总样本数:360,438 条
- 构建方式:由来自四个意大利语语料库(CHANGE-IT、C4-IT、PaISà)的文本,经 14 个 LLM 生成后构建而成。
数据集划分
| 划分 | 样本数 | 说明 |
|---|---|---|
| train | 298,472 | 训练集 |
| validation | 52,672 | 验证集 |
| test | 9,294 | 硬负样本子集——对所有评估系统(AUC 0.39–0.56,接近随机猜测)均具有挑战性 |
用途
用于检测 LLM 生成的意大利语文本,特别强调硬负样本挖掘与基准评测。
引用
该数据集对应的论文正在审核中,拟发表于 CLiC-it 2026(2026年9月14–16日,意大利巴勒莫)。引用格式如下:
bibtex @inproceedings{pisent2026challengeit, title = {CHALLENGE-IT: A Multi-Source Italian Benchmark with Hard Negatives for LLM-Generated Text Detection}, author = {Pisent, Alessandro and Silvestri, Fabrizio}, booktitle = {Proceedings of the Twelfth Italian Conference on Computational Linguistics (CLiC-it 2026)}, year = {2026}, note = {Under review}, address = {Palermo, Italy} }




