Visual_Agent
收藏资源简介:
Visual_Agent 是一个用于视觉工具使用的训练轨迹数据集,其特点在于包含了便携式的图像引用。数据集包含总计 3,679 条训练轨迹,这些轨迹来源于两部分:一部分是 2,187 条基于 DeepEyesV2 衍生的轨迹,涵盖了计数、空间关系和属性相关的任务;另一部分是 1,492 条经过单独评审的 P2R Natural V2 轨迹。所有数据被整合在一个名为 all_training_trajectories_with_images.jsonl 的便携式 JSONL 文件中,其中操作所需的图像路径均相对于 training_trajectories_natural/ 目录。该数据集旨在支持与视觉工具使用相关的机器学习训练任务。
Visual_Agent is a training trajectory dataset for visual tool usage, characterized by including portable image references. The dataset contains a total of 3,679 training trajectories, sourced from two parts: one part consists of 2,187 trajectories derived from DeepEyesV2, covering tasks related to counting, spatial relationships, and attributes; the other part consists of 1,492 individually reviewed P2R Natural V2 trajectories. All data is consolidated in a portable JSONL file named all_training_trajectories_with_images.jsonl, where the image paths required for operations are relative to the training_trajectories_natural/ directory. This dataset aims to support machine learning training tasks related to visual tool usage.
数据集概述
- 数据集名称:Visual_Agent
- 许可证:CC-BY-NC 2.0(非商业用途)
- 数据集来源:Hugging Face(https://huggingface.co/datasets/albert13200/Visual_Agent)
数据集内容
Visual_Agent 是一个包含视觉工具使用训练轨迹的数据集,所有图像引用均为便携式路径格式。训练数据由两个子集组成,总共有 3,679 条轨迹:
训练轨迹来源
- DeepEyesV2 衍生轨迹:共 2,187 条,涵盖计数、空间关系和属性相关的任务轨迹。
- P2R Natural V2 轨迹:共 1,492 条,位于
p2r_v2/子目录下,每条轨迹均经过单独审核。
数据组织
- 所有训练轨迹合并存放于统一文件:
training_trajectories_natural/all_training_trajectories_with_images.jsonl。 - 数据集中所有操作性的图像路径均为相对路径,相对基准目录为
training_trajectories_natural/。
使用说明
数据集包含图像引用,使用时需要相对于 training_trajectories_natural/ 目录解析图像路径。数据集仅限非商业用途(CC-BY-NC 2.0)。





