VideoNet
收藏资源简介:
VideoNet是由华盛顿大学和艾伦人工智能研究所联合构建的大规模领域特定动作识别数据集,涵盖37个领域的1000种精细动作。该数据集包含近50万条视频问答对,视频平均时长12.2秒,数据来源于网络公开视频并经过专业标注流程验证。通过三阶段人工标注流程(视频收集、片段验证和精细修剪)确保数据质量,专家验证显示标签准确率达97%。该数据集旨在推动视觉语言模型在专业领域动作理解方面的研究,可应用于运动分析、医疗动作识别等需要细粒度动作理解的场景。
VideoNet is a large-scale domain-specific action recognition dataset jointly developed by the University of Washington and the Allen Institute for Artificial Intelligence, covering 1000 fine-grained action categories across 37 domains. This dataset contains nearly 500,000 video-question-answer pairs, with an average video duration of 12.2 seconds. The data is sourced from publicly available online videos and validated via a professional annotation workflow. A three-stage manual annotation pipeline consisting of video collection, clip verification and fine-grained trimming is adopted to ensure data quality, and expert validation shows that the label accuracy reaches 97%. This dataset aims to promote research on vision-language models for domain-specific action understanding, and can be applied to scenarios requiring fine-grained action understanding such as motion analysis and medical action recognition.
根据您提供的数据集详情页面内容,以下是数据集的关键信息概述:
- 数据集名称:VideoNet
- 核心内容:包含 1,000 种动作,覆盖 37 个领域。
- 任务/目标:面向特定领域(Domain-Specific)的动作识别,研究背景为视觉语言模型(VLM)时代。
- 相关资源:提供数据(🤗 Data)、演示(Demo)和代码(Code)。
- 会议发表:CVPR 2026 Highlight(亮点论文)。
- 作者单位:University of Washington、Allen Institute for AI、Stanford University。
- 致谢:本项目部分由苹果公司(Apple)资助。




