Human Actions
收藏资源简介:
Human Actions数据集是由日本的Ochanomizu大学创建,专注于捕捉人类动作的动态表达,用于多模态逻辑推理。该数据集包含200个视频,总计1,942个动作标签,每个标签以⟨subject, predicate, object⟩的形式呈现,便于转化为逻辑语义表达。数据集的创建过程涉及视频选择和详细标注,旨在通过复杂的动作描述支持视频与文本间的复杂推理。该数据集的应用领域包括评估视频与复杂语义句子的多模态推理系统,特别是涉及否定和量化表达的情境。
The Human Actions Dataset was developed by Ochanomizu University in Japan, focusing on capturing dynamic expressions of human actions for multimodal logical reasoning. This dataset consists of 200 videos and a total of 1,942 action labels, each presented in the form of ⟨subject, predicate, object⟩ to facilitate conversion into logical semantic expressions. The dataset creation process involves video selection and detailed annotation, aiming to support complex cross-modal reasoning between videos and text through intricate action descriptions. Application scenarios of this dataset include evaluating multimodal reasoning systems for videos and complex semantic sentences, particularly in contexts involving negation and quantificational expressions.

- 1Building a Video-and-Language Dataset with Human Actions for Multimodal Logical InferenceOchanomizu University, Tokyo, Japan · 2021年



