OpenVisTool-42K
收藏资源简介:
OpenVisTool-42K 是一个包含42,048条经过结果验证和工具使用增益过滤的视觉工具使用轨迹数据集。数据集涵盖图表(Chart)、GUI定位(GUI Grounding)、表格(Table)、网页转HTML(Web-to-HTML)和视觉搜索(Visual Search)五个领域。每条轨迹以ms-swift智能体格式保存,包含教师模型的推理过程、函数调用、工具观察结果和最终答案。数据记录为JSONL格式,每个对象包含id(稳定标识符)、domain(领域)、tools(JSON编码的函数模式)、messages(按顺序排列的智能体消息,包括system、user、assistant、tool_call、tool_response角色)和images(相对于快照根目录的图片路径列表,与消息中的<image>标记一一对应)。数据来源包括ChartVerse-SFT-600K、AgentNet、OS-Atlas、UGround、CoSyn-400K、TABLET-Small、DeepEyesV2-RL、Vero-600K和VinciCoder-1.6M-SFT等多个上游数据集,各自适用的许可证延续适用。该数据集适用于多模态工具使用、智能体视觉推理、视觉问答等任务。
OpenVisTool-42K is a dataset of 42,048 visual tool-use trajectories filtered by result verification and tool-use gain. It covers five domains: Chart, GUI Grounding, Table, Web-to-HTML, and Visual Search. Each trajectory is saved in the ms-swift agent format, containing the teacher models reasoning process, function calls, tool observations, and final answers. The data is recorded in JSONL format, with each object including id (stable identifier), domain, tools (JSON-encoded function schema), messages (ordered agent messages with roles system, user, assistant, tool_call, tool_response), and images (list of image paths relative to the snapshot root, corresponding one-to-one with <image> tokens in messages). Data sources include upstream datasets such as ChartVerse-SFT-600K, AgentNet, OS-Atlas, UGround, CoSyn-400K, TABLET-Small, DeepEyesV2-RL, Vero-600K, and VinciCoder-1.6M-SFT, with their respective licenses continuing to apply. The dataset is suitable for tasks such as multimodal tool use, agent visual reasoning, and visual question answering.
OpenVisTool-42K 数据集概述
基本信息
- 数据集名称:OpenVisTool-42K
- 规模:42,048 条轨迹(10K < n < 100K)
- 语言:英语
- 许可证:其他(需遵循上游数据集许可条款)
- 任务类型:图像-文本到文本(image-text-to-text)
- 标签:多模态、工具使用、智能体视觉、视觉推理
数据内容
数据集包含五种领域的视觉工具使用轨迹,每条轨迹保留教师的推理过程、函数调用、工具观察结果和最终答案,格式遵循 ms-swift 智能体格式:
| 领域 | 记录数 |
|---|---|
| 图表 (Chart) | 13,537 |
| GUI 定位 (GUI Grounding) | 11,000+ |
| 表格 (Table) | 4,930 |
| 网页转 HTML (Web-to-HTML) | 10,674 |
| 视觉搜索 (Visual Search) | 1,941 |
数据格式
每条记录为一行 JSON 对象,包含以下字段:
- id:稳定 ID
- domain:所属领域
- tools:JSON 编码的函数模式字符串
- messages:按顺序排列的智能体消息,角色包括 system、user、assistant、tool_call、tool_response
- images:相对于快照根目录的图片路径,与消息中的
<image>标记按顺序对应
数据来源
仅使用源数据集中的图片和查询作为任务输入,不采用预先存在的推理轨迹作为监督。轨迹由教师模型生成,并通过结果有效性和工具使用增益筛选:
| 领域 | 来源数据集 | 许可证 |
|---|---|---|
| 图表 | ChartVerse-SFT-600K | Apache 2.0 |
| GUI 定位 | AgentNet | MIT |
| GUI 定位 | OS-Atlas | Apache 2.0 |
| GUI 定位 | UGround | CC BY-NC-SA 4.0 |
| 表格 | CoSyn-400K | ODC-BY 1.0 |
| 表格 | TABLET-Small | CC BY 4.0 |
| 视觉搜索 | DeepEyesV2-RL | 未指定 |
| 视觉搜索 | Vero-600K | Apache 2.0 |
| 网页转 HTML | VinciCoder-1.6M-SFT | 未指定 |
使用说明
消息中的路径(如 /mnt/data/example.png)为智能体沙箱中的运行时路径,非下载快照路径,训练时不应改写。工具生成的剪裁、掩码、边界框可视化和 HTML 渲染图作为训练轨迹中的观测随原始输入一同打包,其衍生内容受原始图像语料库条款约束。





