vlm_direction
收藏资源简介:
该数据集是一个多配置的视频与文本数据集,适用于视觉问答和视频分类任务。数据集采用MIT许可,但附加了使用限制,禁止用于对人类受试者造成伤害的实验,并强调视频版权归原始创作者或平台所有,仅限学术研究使用。数据集包含12种不同的配置,每种配置对应不同的JSON数据文件,可能代表不同的子任务或数据子集。数据模态包括视频和文本,语言为英语,规模在1,000到15,000条之间。访问数据集需要提供姓名、组织、国家和电子邮件等基本信息。
This dataset is a multi-configuration video-and-text dataset designed for visual question answering (VQA) and video classification tasks. It is released under the MIT License with additional usage restrictions: it is prohibited to use the dataset for experiments that cause harm to human subjects, and it is emphasized that the copyright of the videos belongs to their original creators or platforms, and the dataset is only permitted for academic research purposes. The dataset consists of 12 distinct configurations, each corresponding to a separate JSON data file, which may represent different subtasks or data subsets. Data modalities include video and text, with the text in English, and the dataset size ranges from 1,000 to 15,000 samples. To access the dataset, users are required to provide basic personal and institutional information including name, affiliation, country, and email address.



