ActionHub
收藏资源简介:
ActionHub数据集是一个包含1211种常见人类动作和360万个视频描述的大规模数据集。该数据集通过使用动作名称在视频网站上进行搜索,自动收集视频描述,无需额外的人工标注,因此成本低且易于扩展。ActionHub旨在解决零样本动作识别中的跨模态多样性问题,通过提供丰富的动作视频描述来增强文本模态的语义多样性,从而帮助模型更好地理解视频中的人类动作。数据集的创建过程涉及从七个现有的视频动作数据集中构建动作查询,并通过网络搜索收集相关的视频描述。ActionHub数据集的应用领域主要集中在零样本动作识别,旨在通过学习视频和文本数据之间的有效对齐来识别未见过的动作。
ActionHub Dataset is a large-scale dataset containing 1,211 common human actions and 3.6 million video descriptions. This dataset automatically collects video descriptions by searching video websites with action names, requiring no additional manual annotation, thus featuring low cost and easy scalability. Designed to address the cross-modal diversity issue in zero-shot action recognition, ActionHub enhances the semantic diversity of the text modality by providing abundant action-related video descriptions, thereby enabling models to better comprehend human actions in videos. The dataset construction process involves constructing action queries from seven existing video action datasets and collecting relevant video descriptions via web searches. The primary application scope of the ActionHub Dataset is zero-shot action recognition, where the goal is to recognize unseen actions by learning effective alignment between video and text data.




