ranausmans/feed-injection-pool
收藏资源简介:
该数据集名为Feed-Injection Post Pools,是一个社交媒体风格的帖子池,用于研究推荐系统作为LLM代理的控制表面:对抗性feed注入、模型机制和简单防御。每条帖子以JSON格式存储,包含id、topic、stance、intensity、text等字段,对抗性帖子还包括adversarial、angle、generator字段。数据集分为两部分:远程工作池(headline实验),包括有机帖子、支持返回办公室的对抗性帖子、支持远程工作的对抗性帖子,以及使用Gemma模型生成的变体;泛化任务池(多任务实验),包括安全任务的有机帖子和对抗性帖子,涉及UBI、AI监管、部署安全、供应商安全和访问策略等主题。数据集用于评估推荐系统在LLM代理环境中的对抗性攻击和防御。
The dataset is named Feed-Injection Post Pools, consisting of social-media-style post pools used in the research Recommenders as Control Surfaces for LLM Agents: Adversarial Feed Injection, Model Regimes, and Simple Defenses. Each post is in JSON format with fields such as id, topic, stance, intensity, text, and for adversarial posts, additional fields like adversarial, angle, and generator. It includes two main parts: remote-work pools (headline experiments) with organic posts, adversarial pro-return-to-office and pro-remote posts, and generator-swap variants using Gemma; and generalization-task pools (multi-task experiments) with organic posts for security tasks and adversarial posts pushing attacker-target options across tasks like UBI, AI regulation, deployment security, vendor security, and access policy. The dataset is designed for studying adversarial attacks and defenses in recommender systems within LLM agent contexts.




