PerQ MGTD
收藏资源简介:
PerQ is a dataset (described in a paper) for evaluation of platform-adaptation (Twitter/X, Telegram, and Signal) capabilities of 6 LLMs (smaller and bigger versions of the 3 families: Gemma-3, Qwen3, Llama-3) for 7 languages (English, German, French, Italian, Slovak, Russian, and Hungarian). This version of dataset is intended for augmentation of training sets for machine-generated text detection task. It contains 25,181 machine-generated texts (generated or modified by LLMs using platform-adaptation prompts) and 1,398 human-written headlines and articles of news subsampled from MasiveSumm dataset (about 2x100 texts per language). The dataset has been anonymized to minimize amount of sensitive data by hiding email addresses, usernames, and phone numbers. If you use this dataset in any publication, project, tool or in any other form, please, cite the paper. Disclaimer Due to data source, the dataset may contain harmful, disinformation, or offensive content. Although we have used data sources of older date (lower probability to include machine-generated texts), the labeling (of human-written text) might not be 100% accurate. The anonymization procedure might not successfully hiden all the sensitive/personal content; thus, use the data cautiously (if feeling affected by such content, report the found issues in this regard to dpo[at]kinit.sk). The intended use if for non-commercial research purpose only. Fields The dataset has the following fields: text - a text sample, label - 0 for human-written text, 1 for machine-generated text, multi_label - a string representing a large language model that generated the text or the string "human" representing a human-written text, source - a string identifying the origin of the text, personalization_type - the type of personalization applied in the prompt (Generation/Modification, or "none" in case of original human texts), target_platform - the identifier of the target social-media platform for which the text was adjusted (Twitter, Telegram, Signal, or "news" in case of original human texts), target_language - the intended language of the text (e.g., English, German, etc.), LLM1_personalization_eval - personalization quality label assigned by the first LLM annotator, LLM2_personalization_eval - personalization quality label assigned by the second LLM annotator, LLM3_personalization_eval - personalization quality label assigned by the third LLM annotator, majority_personalization_eval - personalization quality label resulting from inter-LLM majority voting across the three annotators, Basic statistics: label personalization_type target_platform English French German Hungarian Italian Russian Slovak 0 none news 200 200 200 200 200 200 198 1 Generation Signal 600 600 600 599 600 600 599 1 Generation Telegram 600 600 600 600 600 600 600 1 Generation Twitter 600 600 600 600 600 600 600 1 Modification Signal 600 600 600 600 600 600 594 1 Modification Telegram 600 600 600 600 600 600 595 1 Modification Twitter 600 600 600 600 600 600 594



