nemotron-gym-agentic-indirect-prompt-injection
收藏资源简介:
本数据集名为nemotron-gym-agentic-indirect-prompt-injection,由LAION组织提供,源自NVIDIA的Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1数据集,属于Nemotron-Post-Training-v3集合的一部分。数据集包含1,272个任务,每个任务以Harbor任务二进制格式存储,其中每个样本包括path(字符串类型)和task_binary(gzip tar压缩格式)两个字段。数据通过OpenThoughts-Agent框架的data.nemotron_gym工具转换而来,适用于文本生成任务,特别聚焦于代理(agent)和强化学习场景中的间接提示注入(indirect prompt injection)问题。数据集提供了基于单步注入抵抗代理的评分机制:奖励值为1表示代理未发出注入调用。该数据集可用于训练和评估代理在抵抗间接提示注入方面的能力,标签包括agent、harbor、reinforcement-learning和nemotron,采用Apache-2.0许可证。
This dataset is named nemotron-gym-agentic-indirect-prompt-injection. It is provided by LAION, derived from NVIDIA's Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 dataset, and constitutes a part of the Nemotron-Post-Training-v3 collection. The dataset comprises 1,272 tasks, each stored in Harbor task binary format. Each sample contains two fields: "path" (string type) and "task_binary" (gzip tar-compressed format). Converted via the data.nemotron_gym tool from the OpenThoughts-Agent framework, this dataset is tailored for text generation tasks, with a specific focus on indirect prompt injection scenarios in agent and reinforcement learning contexts. It offers a scoring mechanism for single-step injection-resistant agents: a reward value of 1 signifies that the agent did not issue an injection call. This dataset can be utilized to train and evaluate agents' ability to resist indirect prompt injection, with tags including agent, harbor, reinforcement-learning, and nemotron, and it is licensed under the Apache-2.0 license.
数据集概览:laion/nemotron-gym-agentic-indirect-prompt-injection
| 属性 | 内容 |
|---|---|
| 任务类型 | 文本生成(text-generation) |
| 许可证 | Apache-2.0 |
| 样本数量 | 1,272 个任务 |
| 来源 | 由 nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 转换而来,属于 nvidia/Nemotron-Post-Training-v3 系列 |
数据格式
- 每一行是一个有效的 Harbor 任务二进制文件,包含两个字段:
path(字符串)task_binary(gzip tar 压缩格式)
- 转换工具:OpenThoughts-Agent 框架中的
data.nemotron_gym
评估机制
- 评分标准:单步注入抵抗代理(reward 1 表示未执行被注入的调用)




