未明确命名,可称为对抗性AI生成社交机器人内容数据集
收藏资源简介:
本数据集由庞培法布拉大学与印第安纳大学社交网络观察站等机构联合构建,旨在通过对抗性方法模拟恶意行为者在社交媒体上模仿真实用户生成AI内容,为检测AI驱动的社交机器人提供基准数据。数据集包含来自Reddit和Telegram平台的73,521条真实用户消息,通过Gemma-3n-E4B和Qwen3-235B等模型生成对应AI文本,形成263,594对多语言跨平台配对数据,覆盖17种语言和36个频道,平均每条消息约136-156字符。其创建过程采用基于用户历史行为和对话上下文的个性化生成管道,通过严格的数据清洗和语言对齐确保真实性。该数据集主要应用于社交机器人检测、AI生成内容识别及信息安全领域,致力于解决传统检测模型在短文本和跨平台场景下性能不足的问题,以应对虚假信息和身份冒充等威胁。
This dataset was jointly developed by institutions including Pompeu Fabra University and the Indiana University Observatory on Social Media, among others. It is designed to simulate malicious actors generating AI content by mimicking real users on social media platforms through adversarial approaches, and to act as a benchmark dataset for detecting AI-powered social bots. The dataset comprises 73,521 real user messages sourced from Reddit and Telegram platforms. Corresponding AI-generated texts were produced using models such as Gemma-3n-E4B and Qwen3-235B, resulting in 263,594 multilingual cross-platform paired data pairs, covering 17 languages and 36 channels, with an average length of 136 to 156 characters per message. Its development adopts a personalized generation pipeline based on users' historical behaviors and conversational contexts, and guarantees data authenticity via strict data cleaning and language alignment protocols. This dataset is primarily utilized in the fields of social bot detection, AI-generated content recognition, and information security. It aims to mitigate the performance shortcomings of traditional detection models in short-text and cross-platform scenarios, thereby countering threats including disinformation and identity impersonation.




