BARISTA
收藏资源简介:
BARISTA是一个密集标注的以自我为中心(第一人称视角)的视频数据集,专注于咖啡制作场景。该数据集旨在为视觉语言模型(VLMs)在空间理解、时序理解、关系推理和过程理解等多任务上提供一个统一的基准测试平台。数据集包含185个以自我为中心的视频,总时长约4.4小时,帧率为30 FPS,分辨率介于1280×720至1920×1080之间。视频内容涵盖了三种常见的咖啡制作方法:胶囊咖啡机、手柄(portafilter)咖啡机以及全自动咖啡机。所有视频均在受控的室内环境中,使用包括iPhone、Apple Vision Pro、RayBan Meta 3和RayBan Wayfarer智能眼镜在内的多种设备录制。每个视频样本存储于独立目录,包含原始视频文件(video.mp4)和一个扩展的COCO格式标注文件(coco_annotation.json)。该标注文件提供了丰富且结构化的信息:逐帧的对象实例标注(包含边界框和分割掩码)、对象级别的属性(如颜色、状态)、对象之间的有向类型化关系(如位置关系、人机交互动作)、对象类别定义、细粒度的“动词+名词”活动片段、高级别的过程步骤片段、完整的视频元数据(如分辨率、帧数、录制设备)以及数据集划分(训练集/测试集)。BARISTA数据集适用于多种计算机视觉与多模态任务,包括但不限于视频分类、目标检测、视觉问答、活动识别、手物交互分析、实例分割和关系抽取,是评估视觉语言模型组合式理解能力的综合性资源。
BARISTA is a densely annotated egocentric (first-person perspective) video dataset focused on coffee-making scenarios. It aims to provide a unified benchmark for visual language models (VLMs) across multiple tasks such as spatial understanding, temporal understanding, relational reasoning, and procedural understanding. The dataset includes 185 egocentric videos with a total duration of approximately 4.4 hours, a frame rate of 30 FPS, and resolutions ranging from 1280×720 to 1920×1080. The videos cover three common coffee-making methods: capsule coffee machines, portafilter coffee machines, and fully automatic coffee machines. All videos were recorded in controlled indoor environments using various devices including iPhone, Apple Vision Pro, RayBan Meta 3, and RayBan Wayfarer smart glasses. Each video sample is stored in an independent directory, containing the original video file (video.mp4) and an extended COCO-format annotation file (coco_annotation.json). This annotation file provides rich and structured information: per-frame object instance annotations (including bounding boxes and segmentation masks), object-level attributes (e.g., color, state), directed typed relationships between objects (e.g., positional relations, human-machine interaction actions), object category definitions, fine-grained verb+noun activity segments, high-level procedural step segments, comprehensive video metadata (e.g., resolution, frame count, recording device), and dataset splits (train/test). The BARISTA dataset is suitable for various computer vision and multimodal tasks, including but not limited to video classification, object detection, visual question answering, activity recognition, hand-object interaction analysis, instance segmentation, and relation extraction, making it a comprehensive resource for evaluating the compositional understanding capabilities of visual language models.





