VisualTrans
收藏资源简介:
VisualTrans是一个针对现实世界中人与物体交互场景的视觉转换推理(VTR)基准。它包含12种语义丰富的操作任务,并通过系统构建的问题-答案对评估三个核心推理维度——空间、程序和定量。该基准具有472个高质量的问题-答案对,包括选择题、开放式计数和目标枚举等多种格式。VisualTrans基于第一人称操作视频构建,并通过自动元数据注释和结构化问题生成,最终由人工验证确保其高质量和可解释性。该数据集旨在帮助智能系统理解和预测动态场景,并指导行动,为高级智能系统奠定基础。
VisualTrans is a visual transformation reasoning (VTR) benchmark targeting real-world human-object interaction scenarios. It includes 12 semantically rich operational tasks, and evaluates three core reasoning dimensions—spatial, procedural, and quantitative—using systematically constructed question-answer pairs. This benchmark contains 472 high-quality question-answer pairs in diverse formats such as multiple-choice questions, open-ended counting tasks, and object enumeration. VisualTrans is built upon first-person action videos, with automatic metadata annotation and structured question generation applied, followed by manual verification to ensure its high quality and interpretability. This dataset aims to assist intelligent systems in understanding and predicting dynamic scenes, guiding actionable behaviors, and laying a foundation for advanced intelligent systems.
- 1VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning中国科学院自动化研究所 · 2025年



