PerDisNews: Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
收藏资源简介:
The dataset consists of machine-generated personalized disinformation articles and was used to evaluate LLM vulnerabilities to being misused for personalized disinformation generation in a paper accepted to the ACL 2025 Main conference. It consists of 2,268 disinformation articles generated by six LLMs (Vicuna 40B, Falcon 33B, GPT-4o, Llama-3.1-70B, Gemma-2-27b, and Mistral-Nemo). The data were generated using three prompts (without personalization, with personalization by using a simple description of the target group and by using a detailed description of the target group). Generated disinformation articles are about six narratives (health-related and political) and targeted to seven groups (students, parents, seniors, European conservatives, European liberals, rural residents and urban residents). If you use this dataset in any publication, project, tool or in any other form, please, cite the paper. Disclaimer The data are intentional disinformation, generated by large language models. The dataset has been checked for containment of personally identifiable information (PII) and the samples considered dangerous have been anonymized. However, the used procedure might not be 100% effective, therefore, the people or organizations feeling affected can report the found issues in this regard to dpo[at]kinit.sk. Fields The dataset has the following fields: ‘text’ - the text of the generated fake news article ‘target’ - the name of the target group used in the prompt ('-', 'Urban population', 'Seniors', 'European liberals', 'Students', 'European conservatives', 'Parents', 'Rural population’) ‘narrative’ - the title of the disinformation narrative used in the prompt ‘personalization’ - type of prompt used for generation (“no”, “simple”, “detailed”) ‘generator’ - identification of the LLM that generated the text ('Llama-3.1-70B-Instruct', 'Mistral-Nemo-Instruct-2407', 'falcon-40b-instruct', 'gemma-2-27b-it', 'gpt-4o-2024-08-06', 'vicuna-33b-v1.3') ‘annotation_personalization_human1’ - human assigned label of personalization quality on a four-point scale ‘annotation_personalization_human2’ - human assigned label of personalization quality on a four-point scale ‘annotation_personalization_human3’ - human assigned label of personalization quality on a four-point scale ‘annotation_personalization_human4’ - human assigned label of personalization quality on a four-point scale ‘annotation_personalization_human5’ - human assigned label of personalization quality on a four-point scale ‘annotation_personalization_LLM1’ - label of personalization quality assigned by gpt-4o-2024-08-06 meta-evaluator on a four-point scale ‘annotation_personalization_LLM2’ - label of personalization quality assigned by Gemma-2-27b-it meta-evaluator on a four-point scale ‘annotation_personalization_LLM3’ - label of personalization quality assigned by Meta-Llama-3.1-70B-Instruct meta-evaluator on a four-point scale ‘annotation_safetyfilter_LLM2’ - LLM evaluation of presence of safety-filter message in the generated text done by Gemma-2-27b-it (yes/no) 'annotation_noise_LLM2’ - LLM evaluation of presence of noise in the generated text with the narrative done by Gemma-2-27b-it (yes/no) 'annotation_targeted_LLM2’ - LLM evaluation of personalization (targeting) of the generated text with the narrative done by Gemma-2-27b-it (yes/no) ‘annotation_disclaimer_LLM2’ - LLM evaluation of presence of disclaimer message in the generated text done by Gemma-2-27b-it (yes/no) ‘annotation_agreement_LLM2’ - LLM evaluation of agreement of the generated text with the narrative done by Gemma-2-27b-it (yes/no) ‘annotation_disagreement_LLM2’ - LLM evaluation of disagreement of the generated text with the narrative done by Gemma-2-27b-it (yes/no) 'annotation_arguments_against_LLM2' - LLM evaluation of arguments against in the generated text with the narrative done by Gemma-2-27b-it (yes/no) ‘annotation_GRUEN’ - score assigned by automatic metric GRUEN evaluating text quality ‘annotation_LA_LLM2’ - LLM evaluation of generated text quality (linguistic acceptability + output content quality) done by Gemma-2-27b-it ‘annotation_OCQ_LLM2’ - LLM evaluation of generated text quality (output content quality) done by Gemma-2-27b-it ’sentence_count’ - the number of sentences in the generated text



