HindsboNikolaj/scope-benchmark
收藏资源简介:
SCOPE基准测试是一个用于评估自然语言PTZ相机代理的基准数据集,基于HRI 26论文《SCOPE: A Real-Time Natural Language Camera Agent at the Edge》。该数据集仅包含测试集,无训练集,由541个问题、4个Blender场景(包括城市外观、室内房间、城市街道和受损城市外观)和8个任务类别(如计数、描述符、位置/空间、OCR识别、单次调用、多步骤命令、多步骤推理和比较/关系)组成。数据集旨在测试语言模型和视觉模型结合时的性能,并提供详细的评估指标,如问题ID、场景文件路径、自然语言提示、预期答案、评估类别、难度等级等。数据集适用于视觉问答、目标检测和图像分类等任务,重点关注人机交互和机器人视觉语言领域。
The SCOPE benchmark is a benchmark dataset for evaluating natural language-powered PTZ camera agents, based on the HRI 26 paper *SCOPE: A Real-Time Natural Language Camera Agent at the Edge*. This dataset only includes a test split with no training split, and comprises 541 questions, 4 Blender-rendered scenes (covering urban facades, indoor rooms, urban streets, and damaged urban facades), and 8 task categories: counting, descriptor-related queries, location/spatial reasoning, OCR recognition, single-turn invocation, multi-step commands, multi-step reasoning, and comparison/relation tasks. The dataset is designed to evaluate the performance of integrated language and vision models, and provides detailed per-sample evaluation metadata including question ID, scene file path, natural language prompts, expected answers, evaluation category, difficulty level, and more. It is applicable to tasks such as visual question answering, object detection, and image classification, with a core focus on the fields of human-robot interaction and robotic vision-language.




