遇见数据集

AI-TRAITS (AI-generated personalised disinformation dataset)

收藏
Zenodo2026-07-29 更新2026-08-02 收录
官方服务:

资源简介:

Dataset Description The AI-Generated Personalised Disinformation Dataset (AI-TRAITS) is a large-scale, multilingual benchmark resource designed to facilitate research into the vulnerabilities of Large Language Models (LLMs) to generating personalised disinformation. This dataset contains just under 1,6 million texts generated by eight state-of-the-art, instruction-tuned LLMs. The generations are based on 324 unique disinformation narratives and are tailored to 150 distinct demographic personas. The personas are constructed from a combination of four key attributes: country, generation, political orientation, and language. The languages included are English, Russian, Portuguese, and Hindi. The dataset enables a comprehensive analysis of how LLMs adapt disinformation to specific audiences and allows for the evaluation of model safety mechanisms when confronted with persona-targeted prompts. This resource was developed to support the understanding of how simple personalisation strategies can bypass LLM safety features and to highlight the linguistic and rhetorical differences between target-agnostic and persona-targeted AI-generated disinformation. Dataset Fields id: The id of the item in the elastic database; prompt.base_text: The base text of the prompt, without the language or persona instructions; prompt.hash: The hash of the prompt's base text; prompt.generation_iter: The generation iteration of the prompt (1-3); model.family: The family name of the model used to generate the text; output_language: The desired output language for the text; persona.country: The desired country for the text (None for texts with no target persona); persona.generation: The desired generation for the text (None for texts with no target persona); persona.political_orientation: The desired political orientation for the text (None for texts with no target persona); persona.language: The desired language for the text (None for texts with no target persona); persona.hash: The hash for the persona (None for texts with no target persona); country_level: The detected personalisation level for country (None for texts with no target persona, badly generated or detected as refusals); generation_level: The detected personalisation level for generation (None for texts with no target persona, badly generated or detected as refusals); political_orientation_level: The detected personalisation level for political orientation (None for texts with no target persona, badly generated or detected as refusals); response.text: The full response text, including disclaimers, notes, translations, etc.; response.clean_text: The cleaned version of the text without disclaimers, notes, translations, etc.; prompt_entities: The detected entities in the prompt's base text; response_entities: The detected entities in the clean response text; response.model_behaviour: The detected model behaviour (jailbreak, disclaimer, refusal or bad); detected_languages: The detected languages in the response text; persuasions: Persuasion techniques extracted from the clean generated text. Only high precision topics with confidence level above 0.8 selected; response_topics: Topic/news framings extracted from the clean generated text. Only high precision news framings with confidence level above 0.8 selected;prompt_topics: Topic/news framings extracted from the input prompt text. Only high precision news framings with confidence level above 0.8 selected; Ethical Disclaimer This dataset was created for research purposes to investigate the capabilities of Large Language Models (LLMs) in generating persona-targeted disinformation, with the primary objective of identifying system vulnerabilities and contributing to the development of safer AI. The research involved the deliberate generation of false and potentially harmful content in a controlled and secure environment. No generated content has been disseminated publicly, and the study did not involve human subjects or the use of any personal data. All personas used for generation are synthetic. To prevent potential misuse, the AI-TRAITS dataset is made available to researchers upon request, for purposes that align with AI safety, transparency, and responsible technology development. All users must agree to terms that prohibit the redistribution and misuse of the data. The authors' intent is to support harm prevention and improve the safety mechanisms in large-scale generative models. Supplementary Files personalisation_extraction_prompt.txt: Contains the prompt employed to extract personalisation labels. personalisation_prompt.txt: Contains the instruction prompt employed to condition the text generation on a given persona. annotated_test_set.csv: A sample of the dataset that is annotated for personalisation and model behaviour.

提供机构:
Zenodo
创建时间:
2026-07-29
二维码
社区交流群
二维码
科研交流群
商业服务