EgoXR-GUI
收藏资源简介:
EgoXR-GUI是首个专门针对扩展现实(XR)的图形用户界面(GUI)定位基准测试数据集。与传统的桌面或移动GUI基准测试不同,该数据集旨在评估多模态大语言模型(MLLMs)在混合数字-物理环境中对嵌入式虚拟界面的推理能力。数据集包含1,070个精心策划的示例,这些示例采集自Apple Vision Pro及其他3D/XR环境。每个样本包括以下字段:唯一标识符(task_id、annotation_id、sample_id)、XR头戴设备捕获的第一人称视角图像(image)、英文和中文的定位指令(instruction_en、instruction_cn)、表示用户注意力的眼动追踪坐标(gaze_point)、包含上下文标签的结构化字典(choices,如is_same_window、ui_type、platform等)、目标边界框的几何信息(target_bbox,包含坐标、尺寸、旋转和标签)、用于数据查看器可视化的标准化边界框格式(objects)以及质量控制布尔指标(is_ok)。数据集支持三种任务类型:直接定位(简单识别)、空间定位(基于3D空间属性的UI元素推理)和语义定位(基于UI元素的文本或图标语义推理)。适用于视觉定位和物体检测任务,属于计算机视觉、多模态和第一人称视角(XR/GUI)研究领域。
EgoXR-GUI is the first dedicated benchmark dataset for graphical user interface (GUI) localization in extended reality (XR). Unlike traditional desktop or mobile GUI benchmarks, this dataset is designed to evaluate the reasoning abilities of multimodal large language models (MLLMs) toward embedded virtual interfaces in hybrid digital-physical environments. The dataset comprises 1,070 meticulously curated examples collected from Apple Vision Pro and other 3D/XR environments. Each sample contains the following fields: unique identifiers (task_id, annotation_id, sample_id), first-person perspective images captured by XR head-mounted devices (image), localization instructions in English and Chinese (instruction_en, instruction_cn), eye-tracking coordinates representing user attentional focus (gaze_point), a structured dictionary with contextual labels (choices, including is_same_window, ui_type, platform, etc.), geometric information of the target bounding box (target_bbox, encompassing coordinates, dimensions, rotation and label), standardized bounding box format for data viewer visualization (objects), and a quality control boolean metric (is_ok). The dataset supports three task categories: direct localization (simple recognition), spatial localization (UI element reasoning based on 3D spatial attributes), and semantic localization (semantic reasoning of UI elements based on their text or iconography). It is suitable for visual localization and object detection tasks, and falls within the research domains of computer vision, multimodality, and first-person perspective (XR/GUI).
EgoXR-GUI:物理-数字扩展现实中的图形用户界面基准测试
EgoXR-GUI 是首个专为扩展现实(XR)环境设计的图形用户界面(GUI)定位基准数据集。它主要用于评估多模态大语言模型(MLLMs)在混合数字-物理环境中对虚拟界面的推理能力。
核心特征
- 数据集规模:包含 1,070 个经过精心挑选的样本。
- 硬件平台:基于 Apple Vision Pro 及其他 3D/XR 设备采集。
- 任务类型:
- 直接定位:简单界面元素识别。
- 空间定位:基于 UI 元素的 3D 空间属性进行推理。
- 语义定位:基于 UI 元素的文本或图标语义进行推理。
- 支持语言:英文 (
instruction_en) 和中文 (instruction_cn)。
数据字段
每条样本包含以下字段:
task_id与annotation_id:用于追踪特定视觉任务的唯一标识符。sample_id:关联回原始数据集源的外部样本标识符。image:从 XR 头显/环境中捕获的自我中心视角图像。instruction_en/instruction_cn:英文和中文的定位提示指令。gaze_point:表示用户注意力的眼动追踪坐标[x, y]。choices:结构化字典,包含上下文标签(如is_same_window、ui_type、platform、scenario、place、activity、task type)。target_bbox:精确的几何目标,包含坐标、尺寸、空间旋转角度和标签。objects:以 Hugging Face 标准格式表示的边界框,用于数据可视化。is_ok:质量控制布尔值指示器。
许可证与标签
- 许可证:cc-by-4.0
- 任务类别:视觉定位 (visual-grounding)、物体检测 (object-detection)
- 相关标签:computer-vision, visual-grounding, xr, egocentric, gui, apple-vision-pro, instruction-following, multimodal





