HuggingFaceM4/something_something_v2
收藏资源简介:
Something-Something数据集(版本2)是一个包含220,847个标记视频片段的集合,展示了人类使用日常物品执行预定义的基本动作。该数据集旨在训练机器学习模型,以精细理解人类手势,如将某物放入某物、将某物倒置和用某物覆盖某物。数据集支持的任务是动作识别,目标是对视频中发生的动作进行分类。数据集的注释语言为英语,数据实例包括视频ID、视频文件、文本描述、标签和占位符。数据集分为训练集、验证集和测试集,分别包含168,913、24,777和27,157个样本。数据集的创建目的是通过视频预测任务来增强对物理世界的常识理解。数据集的来源是通过众包工人根据给定的标签收集视频。
The Something-Something Dataset (Version 2) is a collection of 220,847 annotated video clips that show humans performing predefined basic actions with everyday objects. This dataset aims to train machine learning models to achieve fine-grained understanding of human gestures, such as putting an object into another, inverting an object, and covering an object with another. The task supported by this dataset is action recognition, whose goal is to classify the actions occurring in a video. The annotation language of this dataset is English, and each data instance includes video ID, video file, text description, label, and placeholder. The dataset is split into training, validation, and test sets, containing 168,913, 24,777, and 27,157 samples respectively. The dataset was created to enhance common-sense understanding of the physical world via video prediction tasks. The dataset was compiled by collecting videos from crowdworkers based on given labels.
数据集概述
数据集名称
- 名称: Something Something v2
- 别名: Something-Something dataset (version 2)
数据集描述
- 摘要: Something Something v2 是一个包含220,847个标记视频片段的数据集,这些视频展示了人类执行预定义的基本动作,使用日常物品。该数据集旨在训练机器学习模型,以理解精细的人类手势,如将某物放入某物中,将某物倒置,以及用某物覆盖某物。
- 语言: 数据集的标注语言为英语。
数据集结构
- 数据实例: 每个数据实例包含视频ID、视频文件、文本描述、标签和占位符。
- 数据字段:
video_id: 视频的唯一标识符。video: 视频文件对象。placeholders: 视频中出现的对象列表。text: 视频中发生的事件描述。labels: 视频中的动作标签,范围从0到173。
数据集创建
- 来源: 数据集为原创数据,由众包工作者提供视频和标签。
- 标注过程: 标签先于视频收集,由AMT工作者完成。
数据集使用考虑
- 社会影响: 该数据集对于动作识别预训练非常有用,因其包含多样化的动作。
- 许可证: 数据集的许可证为QualComm定义的一页文档,使用前需详细阅读。
引用信息
bibtex @inproceedings{goyal2017something, title={The" something something" video database for learning and evaluating visual common sense}, author={Goyal, Raghav and Ebrahimi Kahou, Samira and Michalski, Vincent and Materzynska, Joanna and Westphal, Susanne and Kim, Heuna and Haenel, Valentin and Fruend, Ingo and Yianilos, Peter and Mueller-Freitag, Moritz and others}, booktitle={Proceedings of the IEEE international conference on computer vision}, pages={5842--5850}, year={2017} }




