Jaehwisong/XVR
收藏资源简介:
XVR(Cross-View Relations)是一个大规模数据集,旨在教授视觉语言模型(VLMs)跨多个视角的空间推理能力。它包含来自18K个多样化3D场景(通用领域)和70K个机器人操作轨迹(机器人领域)的100K个视觉问答样本,涵盖三个基本空间推理任务:对应性、验证和定位。通过在XVR上微调,VLMs在多视角和机器人空间推理基准测试(如MindCube、RoboSpatial)上取得明显改进,并且当用作视觉语言动作(VLA)模型的主干时,能平均提高RoboCasa上的操作成功率约13%。
XVR (Cross-View Relations) is a large-scale dataset designed to teach Vision-Language Models (VLMs) spatial reasoning across multiple viewpoints. It comprises 100K vision-question-answer samples derived from 18K diverse 3D scenes (general domain) and 70K robotic manipulation trajectories (robotic domain), spanning three fundamental spatial reasoning tasks: Correspondence, Verification, and Localization. VLMs fine-tuned on XVR achieve clear improvements on multi-view and robotic spatial reasoning benchmarks (MindCube, RoboSpatial), and when used as backbones in Vision-Language-Action (VLA) models, improve manipulation success rates on RoboCasa by ~13% absolute on average.




