CATER
收藏资源简介:
CATER数据集由卡内基梅隆大学创建,专注于视频中的组合动作和时间推理分析。该数据集包含5500个合成渲染的视频,每个视频长度为10秒,使用标准的3D对象库生成,旨在测试模型对长期时间推理的能力。CATER不仅是一个具有挑战性的数据集,还提供了丰富的诊断工具,用于分析现代视频架构。数据集的应用领域包括动作识别、组合动作识别和目标跟踪等,旨在解决视频理解中的复杂时间推理问题。
The CATER dataset was created by Carnegie Mellon University, focusing on compositional action and temporal reasoning analysis in videos. This dataset contains 5,500 synthetically rendered videos, each with a duration of 10 seconds, generated using standard 3D object libraries, and is designed to test models' capabilities for long-term temporal reasoning. Beyond being a challenging dataset, CATER also provides rich diagnostic tools for analyzing modern video architectures. The application areas of the dataset include action recognition, compositional action recognition, object tracking and other fields, aiming to solve complex temporal reasoning problems in video understanding.

- 1CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning卡内基梅隆大学 · 2020年



