遇见数据集

CEAID

收藏
Zenodo2026-05-26 更新2026-05-29 收录
官方服务:

资源简介:

CEAID is a dataset (described in a paper) for machine-generated text detection benchmark for 7 Central European languages (Croatian, Czech, German, Hungarian, Polish, Slovak, and Slovenian) in two domains (news and social media). It contains 188,098 texts, of which about 23k are human-written and about 165k are generated by 8 multilingual large language models. The dataset has been anonymized to minimize amount of sensitive data by hiding email addresses, usernames, and phone numbers. If you use this dataset in any publication, project, tool or in any other form, please, cite the paper. Disclaimer Due to data source, the dataset may contain harmful, disinformation, or offensive content. MultiSocial dataset description states that based on a multilingual toxicity detector, about 8% of the text samples are probably toxic (from 5% in WhatsApp to 10% in Twitter). Although we have used data sources of older date (lower probability to include machine-generated texts), the labeling (of human-written text) might not be 100% accurate. The anonymization procedure might not successfully hiden all the sensitive/personal content; thus, use the data cautiously (if feeling affected by such content, report the found issues in this regard to dpo[at]kinit.sk). The intended use if for non-commercial research purpose only. Data Source The dataset is a subset of a combination of data from MULTITuDEv3 (news articles) and MultiSocial (social-media texts). The data contain at least 200 test samples per each domain and class (human vs. machine) for each of the selected Central European languages (Croatian, Czech, German, Hungarian, Polish, Slovak, and Slovenian). The machine texts are generated by 8 LLMs, 6 of which are the same across the two domains (Aya-101, GPT-3.5-Turbo-0125, Mistral-7B-Instruct-v0.2, OPT-IML-Max-30B, v5-Eagle-7B-HF, and Vicuna-13B), one is only in news domain (Llama-2-70B-chat-hf), and one is only in social-media domain (Gemini). The dataset has the following fields: 'text' - a text sample, 'label' - 0 for human-written text, 1 for machine-generated text, 'multi_label' - a string representing a large language model that generated the text or the string "human" representing a human-written text, 'split' - a string identifying train or test split of the dataset for the purpose of training and evaluation respectively, 'language' - the ISO 639-1 language code identifying the detected language of the given text, 'length' - word count of the given text, 'source' - a string identifying the source dataset / platform of the given text, 'domain' - "news" for news articles from MULTITuDE, "social_media" for social-media texts from MultiSocial. Basic statistics: Language News Train News Test Social Media Train Social Media Test All Train All Test cs (Czech) 7734 2328 11041 6073 18775 8401 de (German) 7764 2322 21038 9497 28802 11819 hr (Croatian) 7819 2348 14475 5993 22294 8341 hu (Hungarian) 7791 2350 14492 5957 22283 8307 pl (Polish) 7818 2336 16687 6971 24505 9307 sk (Slovak) 7664 2317 0 2026 7664 4343 sl (Slovenian) 7845 2354 0 3058 7845 5412 Total 54435 16355 77733 39575 132168 55930

提供机构:
Zenodo
创建时间:
2026-05-26
二维码
社区交流群
二维码
科研交流群
商业服务