ActionArt
收藏资源简介:
ActionArt是一个细粒度视频字幕数据集,旨在推动以人为中心的多模态理解研究。该数据集包含数千个视频,捕捉了广泛的人类动作、人-物交互和多样化的场景,每个视频都伴有详细的注释,精确地标注了每个肢体动作。我们开发了八个子任务,以评估现有大型多模态模型在不同维度上的细粒度理解能力。实验结果表明,尽管当前的大型多模态模型在各种任务上表现良好,但它们往往在实现细粒度理解方面有所欠缺。我们认为,这种局限性主要归因于精细标注数据的稀缺,这些数据既昂贵又难以手动扩展。由于手动注释既昂贵又难以扩展,我们提出了代理任务来增强模型在空间和时间维度上的感知能力。这些代理任务经过精心设计,由现有大型语言模型自动生成的数据驱动,从而减少了对外部昂贵手动标签的依赖。实验结果表明,提出的代理任务显著缩小了与手动标注细粒度数据相比的性能差距。
ActionArt is a fine-grained video captioning dataset designed to advance human-centric multimodal understanding research. This dataset contains thousands of videos that capture a wide range of human actions, human-object interactions and diverse scenarios, with each video accompanied by detailed annotations that precisely label every limb movement. We developed eight subtasks to evaluate the fine-grained understanding capabilities of state-of-the-art large multimodal models across diverse dimensions. Experimental results show that although current large multimodal models perform well across various tasks, they often fall short in achieving fine-grained understanding. We argue that this limitation is mainly attributed to the scarcity of finely annotated data, which is both costly and difficult to scale manually. Given that manual annotation is both costly and difficult to scale, we propose proxy tasks to enhance the perceptual capabilities of models across spatial and temporal dimensions. These proxy tasks are carefully designed and data-driven, automatically generated by existing large language models, thereby reducing reliance on expensive external manual annotations. Experimental results demonstrate that the proposed proxy tasks significantly narrow the performance gap compared to using manually annotated fine-grained data.
数据集概述
基本信息
- 数据集名称: maybex/ActionArt
- 许可证: Apache License 2.0
- 创建者: @maybex
- 下载量: 0
- 大小: 5.18MB
- 更新时间: 2025-04-28
数据集状态
- 当前状态: 内容尚未更新,请期待后续更新。
备注
- 该数据集由ModelScope.cn平台提供。

- 1ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding中山大学, 中国; 阿里巴巴集团通义实验室; 深圳鹏城实验室, 中国; 教育部机器智能与先进计算重点实验室, 中国 · 2025年



