DECADE
收藏资源简介:
我们介绍了直接对视觉智能代理进行建模的任务。计算机视觉通常专注于解决与视觉智能相关的各种子任务。我们偏离了计算机视觉的这种标准方法;相反,我们直接对视觉智能代理进行建模。我们的模型将视觉信息作为输入并直接预测代理的动作。为此,我们引入了 DECADE,这是一个从狗的角度以及她相应的动作的以自我为中心的视频的大规模数据集。使用这些数据,我们模拟了狗的行为以及狗如何计划她的动作。我们在各种指标下表明,仅给定视觉输入,我们就可以在许多情况下成功地对这个智能代理进行建模。此外,与图像分类训练的表示相比,我们模型学习的表示编码了不同的信息,并且我们学习的表示可以推广到其他领域。特别是,我们通过使用这种狗建模任务作为表示学习,在可步行表面估计任务上展示了强大的结果。
We introduce the task of directly modeling visual intelligent agents. Computer vision typically focuses on solving various subtasks related to visual intelligence. We depart from this standard approach in computer vision; instead, we directly model visual intelligent agents. Our model takes visual information as input and directly predicts the agent's actions. To this end, we introduce DECADE, a large-scale dataset of egocentric videos from a dog's perspective paired with her corresponding actions. Using this data, we model the dog's behavior and how she plans her movements. We demonstrate across various metrics that given only visual inputs, we can successfully model this intelligent agent in many scenarios. Furthermore, the representations learned by our model encode distinct information compared to those trained for image classification, and our learned representations can generalize to other domains. Specifically, we show strong results on the traversable surface estimation task by using this dog modeling task as a form of representation learning.




