EPIC-KITCHENS VISOR
收藏资源简介:
We introduce VISOR, a new dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video. VISOR annotates videos from EPIC-KITCHENS, which comes with a new set of challenges not encountered in current video segmentation datasets. Specifically, we need to ensure both short- and long-term consistency of pixel-level annotations as objects undergo transformative interactions, e.g. an onion is peeled, diced and cooked - where we aim to obtain accurate pixel-level annotations of the peel, onion pieces, chopping board, knife, pan, as well as the acting hands. VISOR introduces an annotation pipeline, AI-powered in parts, for scalability and quality. Data published under the Creative Commons Attribution-NonCommerial 4.0 International License.
本研究推出VISOR数据集(VISOR),其包含像素级标注信息,同时配套面向第一人称视角视频中手部与活动物体分割任务的基准测试套件。VISOR的标注视频源自EPIC-KITCHENS数据集,该数据集所面临的一系列全新挑战为现有视频分割数据集所未见。具体而言,当物体经历各类转化性交互操作(例如对洋葱进行剥皮、切丁与烹饪)时,需确保像素级标注的短期与长期一致性;本数据集旨在精准获取洋葱外皮、洋葱块、砧板、刀具、平底锅以及操作手部的像素级标注信息。VISOR配套开发了兼具可扩展性与标注质量的标注流程,该流程部分采用人工智能技术赋能。本数据集依据知识共享署名-非商业性使用4.0国际许可协议(Creative Commons Attribution-NonCommercial 4.0 International License)发布。




