遇见数据集

Nemotron-RL-Instruction-Following-Adversarial-v1

收藏
魔搭社区2026-07-14 更新2026-07-15 收录
官方服务:

资源简介:

## Dataset Description: The inverseIF dataset focuses on adversarial prompts designed to explicitly conflict with an AI model’s standard training instincts—such as writing code without comments or refusing standard helpfulness norms—across 8 distinct "anti-convention" patterns. Using a targeted "model breaking" methodology, it generates four candidate responses via Nemotron-Nano-V2 or Qwen3-235B-A22B-Thinking-2507 to test if the negative constraint is difficult enough to force a default-behavior failure. These responses are rigorously evaluated by both human judges and a GPT-5 LLM judge (requiring at least an 85% agreement rate between them), and a sample is only accepted if it successfully "breaks" the models—meaning at most one of the four responses passes the strict rubric while still demonstrating variance (at least one pass and one fail). The resulting challenging samples are formatted into comprehensive JSON files containing the adversarial prompt, ground truth, candidate responses, and the detailed dual-evaluation metrics. This dataset is released as part of NVIDIA [NeMo Gym](https://github.com/NVIDIA-NeMo/Gym), a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training environments and datasets to enable Reinforcement Learning from Verifiable Reward (RLVR). This dataset was utilized in the development of the [NVIDIA Nemotron](https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/) family of models. NeMo Gym is an open-source library within the [NVIDIA NeMo framework](https://github.com/NVIDIA-NeMo/), NVIDIA's GPU accelerated, end-to-end training framework for large language models (LLMs), multi-modal models and speech models. This dataset is part of the https://huggingface.co/collections/nvidia/nemo-gym/ collection This dataset is ready for commercial use. ## Dataset Owner(s): NVIDIA Corporation ## Dataset Creation Date: 03/11/2026 ## License/Terms of Use: This dataset is licensed under Creative Commons Attribution 4.0 International (CC-BY 4.0). ## Intended Usage: To be used with [NeMo Gym](https://github.com/NVIDIA-NeMo/Gym) for post-training LLMs. ## Dataset Characterization * Data Collection Method<br> * Hybrid: Synthetic, Human * Labeling Method<br> * Hybrid: Synthetic, Human ## Dataset Format Structured JSON, Compatible with https://huggingface.co/collections/nvidia/nemo-gym ## Dataset Quantification [Record Count- 100 entries]<br> [Feature Count- 4 top layers]<br> [Measurement of Total Data Storage- 72MB] ## Reference(s): [NeMo Gym](https://github.com/NVIDIA-NeMo/Gym)<br> [Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?](https://arxiv.org/abs/2509.04292) ## Ethical Considerations: NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse. Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).

提供机构:
maas
创建时间:
2026-03-12
二维码
社区交流群
二维码
科研交流群
商业服务