遇见数据集

MULTITuDE

收藏
Zenodo2025-05-26 更新2026-05-26 收录
官方服务:

资源简介:

MULTITuDE is a dataset for multilingual machine-generated text detection benchmark, described in the EMNLP 2023 conference paper. It consists of 7992 human-written news texts in 11 languages subsampled from MassiveSumm, accompanied by 66089 texts generated by 8 large language models (by using headlines of news articles). The creation process and scripts for replication/extension are located in a GitHub repository. If you use this dataset in any publication, project, tool or in any other form, please, cite the paper. Fields The dataset has the following fields: 'text' - a text sample, 'label' - 0 for human-written text, 1 for machine-generated text, 'multi_label' - a string representing a large language model that generated the text or the string "human" representing a human-written text, 'split' - a string identifying train or test split of the dataset for the purpose of training and evaluation respectively, 'language' - the ISO 639-1 language code identifying the language of the given text, 'length' - word count of the given text, 'source' - a string identifying the source dataset / news medium of the given text. Statistics (the number of samples) Splits: train - 44786 test - 29295 Binary labels: 0 - 7992 1 - 66089 Multiclass labels: gpt-3.5-turbo - 8300 gpt-4 - 8300 text-davinci-003 - 8297 alpaca-lora-30b - 8290 vicuna-13b - 8287 opt-66b - 8229 llama-65b - 8229 opt-iml-max-1.3b - 8157 human - 7992 Languages: English (en) - 29460 (train + test) Spanish (es) - 11586 (train + test) Russian (ru) - 11578 (train + test) Dutch (nl) - 2695 (test) Catalan (ca) - 2691 (test) Czech (cs) - 2689 (test) German (de) - 2685 (test) Chinese (zh) - 2683 (test) Portuguese (pt) - 2673 (test) Arabic (ar) - 2673 (test) Ukrainian (uk) - 2668 (test)

MULTITuDE是一款面向多语言机器生成文本检测的基准数据集,相关研究成果发表于EMNLP 2023会议论文。该数据集从MassiveSumm中采样得到11种语言共7992条人类撰写的新闻文本,同时配套了由8个大语言模型(Large Language Model)基于新闻文章标题生成的66089条文本。该数据集的创建流程与用于复现、扩展该数据集的代码脚本均托管于GitHub仓库中。 若在任何出版物、项目、工具或其他形式的应用中使用该数据集,请引用该论文。 字段说明 该数据集包含以下字段: `text` - 文本样本 `label` - 分类标签,0代表人类撰写的文本,1代表机器生成的文本 `multi_label` - 字符串类型字段,若为机器生成文本则标识生成该文本的大语言模型,若为人类撰写文本则取值为字符串"human" `split` - 字符串类型字段,用于标识数据集的训练划分(train)与测试划分(test),分别用于模型训练与评估 `language` - ISO 639-1标准语言代码,用于标识当前文本的语言 `length` - 当前文本的单词总数 `source` - 字符串类型字段,用于标识当前文本的来源数据集或新闻媒体 样本统计(样本数量) 数据集划分: train - 44786 test - 29295 二分类标签分布: 0(人类撰写) - 7992 1(机器生成) - 66089 多分类标签分布: gpt-3.5-turbo - 8300 gpt-4 - 8300 text-davinci-003 - 8297 alpaca-lora-30b - 8290 vicuna-13b - 8287 opt-66b - 8229 llama-65b - 8229 opt-iml-max-1.3b - 8157 human(人类撰写) - 7992 语言分布: 英语(en) - 29460(训练集+测试集) 西班牙语(es) - 11586(训练集+测试集) 俄语(ru) - 11578(训练集+测试集) 荷兰语(nl) - 2695(仅测试集) 加泰罗尼亚语(ca) - 2691(仅测试集) 捷克语(cs) - 2689(仅测试集) 德语(de) - 2685(仅测试集) 中文(zh) - 2683(仅测试集) 葡萄牙语(pt) - 2673(仅测试集) 阿拉伯语(ar) - 2673(仅测试集) 乌克兰语(uk) - 2668(仅测试集)

提供机构:
Zenodo
创建时间:
2023-10-17
二维码
社区交流群
二维码
科研交流群
商业服务