xlangai/RoboFine-bench
收藏资源简介:
RoboFine-Bench是一个用于评估视觉语言模型(VLMs)是否能够捕捉机器人操作执行级细节的基准测试,它超越了粗粒度任务识别,专注于理解机器人如何执行任务。该基准是FineVLA框架的一部分,用于视觉语言动作学习中的细粒度指令对齐。它包含来自10个机器人数据集的500个保留的机器人操作视频,覆盖32种实施例、多样化的相机视角和广泛的操作任务。每个轨迹都配有人工审核的步骤级注释,分解为10,816个原子事实,跨越十个动作相关维度,平均每个样本有4.3个步骤和21.6个事实。所有500个基准轨迹与RoboFine-VLM SFT训练集和所有策略训练分割严格不相交,确保零数据泄露。评估包括两个轨道:VQA轨道(通过1,030个问题评估判别性理解)和Caption轨道(通过生成步骤级描述评估生成性理解)。
RoboFine-Bench is a benchmark for evaluating whether Vision-Language Models (VLMs) can capture execution-level details of robot manipulation — going beyond coarse task recognition to understand how a robot performs a task. It is part of the FineVLA framework for fine-grained instruction alignment in Vision-Language-Action learning. The benchmark contains 500 held-out robot manipulation videos from 10 robot datasets, covering 32 embodiments, diverse camera views, and a wide range of manipulation tasks. Each trajectory is paired with human-reviewed step-level annotations decomposed into 10,816 atomic facts across ten action-relevant dimensions, with an average of 4.3 steps and 21.6 facts per sample. All 500 benchmark trajectories are strictly disjoint from both the RoboFine-VLM SFT training set and all policy-training splits, ensuring zero data leakage. Evaluation includes two tracks: the VQA track (evaluating discriminative understanding through 1,030 questions) and the Caption track (evaluating generative understanding by producing step-level descriptions).




