遇见数据集

RogueGPT Stimulus Corpus: A Multilingual LLM-Generated News Dataset

收藏
Zenodo2026-08-12 更新2026-08-13 收录
官方服务:

资源简介:

A multilingual corpus of 3,278 text fragments generated by large language models and sourced from human journalists, designed for research on AI-generated disinformation detection. The corpus follows a four-quadrant design crossing content origin (Human vs. Machine) with veracity (Real vs. Fake) across four languages (English, German, French, Spanish) and three formats (tweet, headline, short article). It contains 2,638 machine-generated and 640 human-authored fragments from 10 model identifiers across 6 providers, including GPT-3.5 Turbo, GPT-4o, GPT-4.1, Gemma 7B, Gemini 3 Flash, Llama 2 13B, Mistral 7B, Phi-3 Mini and Claude Opus 4.6. Snapshot of 23 March 2026. This version supersedes the initial release of 2,308 fragments. Note for machine learning use: the corpus contains duplicate content under distinct fragment identifiers, so group by normalized text before creating train and test splits. See the codebook for details. Created using the RogueGPT platform as part of a doctoral dissertation on the AI-driven disinformation ecosystem at Frankfurt University of Applied Sciences.

提供机构:
Zenodo
创建时间:
2026-08-12
二维码
社区交流群
二维码
科研交流群
商业服务