ARM-Thinker-Data
收藏资源简介:
ARM-Thinker-Data 是一个用于训练 ARM-Thinker(一种基于工具使用和视觉接地的多模态奖励模型)的数据集。该数据集由 Qwen3-VL-235B-A22B-Instruct、Qwen3-VL-235B-A22B-Thinking 和 GPT-4o 标注,数据文件位于 `qwen/` 目录下。数据集支持多模态任务,包括工具使用、视觉理解和多步推理,涵盖图像裁剪、文档检索、OCR 等多种工具类型。数据分为 SFT(监督微调)和 RL(强化学习)两部分,每个样本通常包含查询、图像、多轮交互轨迹(思考过程、工具调用、观察结果、最终答案)和奖励信号。数据集采用 CC BY-NC 4.0 许可,仅限研究使用。
ARM-Thinker-Data is a dataset designed for training ARM-Thinker, a multimodal reward model based on tool use and visual grounding. This dataset is annotated by Qwen3-VL-235B-A22B-Instruct, Qwen3-VL-235B-A22B-Thinking, and GPT-4o, with its data files stored in the `qwen/` directory. It supports multimodal tasks including tool use, visual understanding and multi-step reasoning, covering various tool types such as image cropping, document retrieval, OCR and more. The dataset is split into two subsets: SFT (Supervised Fine-Tuning) and RL (Reinforcement Learning). Each sample typically contains a query, images, multi-turn interaction trajectories (including thinking process, tool calls, observation results and final answers), as well as reward signals. This dataset is licensed under CC BY-NC 4.0 and is restricted to research use only.




