MicroAGI-Labs/vlm-info-loss-results
收藏资源简介:
该数据集用于评估视觉语言模型(VLM)在机器人操作数据集上的空间信息保留能力。研究背景显示,VLM连接器对视觉表示进行了非线性转换,以优化LLM的输入。评估协议包括两轮空间信息测试,涉及8个机器人操作数据集和3个摄像头视角。数据集包含多个模型家族的评估结果,如Gemma4和Qwen3.5,并提供了不同的评估变体,如标准评估和注意力图分析。关键发现包括Gemma4在生成有效边界框方面的优势,以及手腕摄像头对空间信息保留的负面影响。
This dataset is designed to evaluate the spatial information retention capability of Vision-Language Models (VLMs) on robotic manipulation datasets. Research background indicates that VLM connectors perform non-linear transformations on visual representations to optimize the input for Large Language Models (LLMs). The evaluation protocol consists of two rounds of spatial information tests, involving 8 robotic manipulation datasets and 3 camera viewpoints. The dataset includes evaluation results from multiple model families, such as Gemma4 and Qwen3.5, and provides various evaluation variants, including standard evaluation and attention map analysis. Key findings include the advantage of Gemma4 in generating valid bounding boxes, as well as the negative impact of wrist-mounted cameras on spatial information retention.




