Automatic Understanding of Image and Video Advertisements
收藏资源简介:
视频数据:2,003个广告视频,时长从30秒到2分30秒不等,总时长为1,749分钟。注释:68个视频级标签,重点关注主题和情感。我们通过人工注释为相同的视频数据增加了额外的摘要。增强注释包括:专注于长篇编辑描述的视频级摘要,包括故事情节、意图、信息、语调和目标受众;100个视频级标签,包括类型、格式、主题、情绪和主题(注意:计划在2026年第二季度的下一个MediaPerf版本中发布)。
Video Dataset: 2,003 advertising videos, with durations ranging from 30 seconds to 2 minutes and 30 seconds, with an aggregate duration of 1,749 minutes. Annotations: 68 video-level labels focusing on topics and sentiments. We added supplementary summaries via manual annotations to this identical video dataset. Enhanced annotations include: video-level summaries centered on long-form editorial descriptions, covering storylines, intentions, core information, tones, and target audiences; 100 video-level labels including genres, formats, topics, emotions, and themes (Note: Planned to be released in the next MediaPerf version in Q2 2026).
MediaPerf 数据集概述
数据集基本信息
- 数据集名称:MediaPerf 基准测试数据集
- 核心数据来源:基于“Automatic Understanding of Image and Video Advertisements”研究的视频数据与标注
- 主要用途:用于评估多模态基础模型在视频理解任务上的性能,任务基于媒体行业实际生产环境中的真实数据和需求。
数据内容
- 视频数据:包含 2,003 条广告视频,视频长度范围从 30 秒到 2 分 30 秒,总时长为 1,749 分钟。
- 原始标注:包含 68 个视频级标签,侧重于主题和情感。
- 增强标注:
- 视频级摘要:包含人工标注的长篇编辑性描述,涵盖故事情节、意图、信息、语气和目标受众。
- 视频级标签:包含 100 个标签,涵盖类型、格式、主题、情绪和主题(计划在 2026 年第二季度发布的下一版 MediaPerf 中提供)。
数据文件与格式
- 视频ID列表:位于
data/inputs/youtube_video_ids.txt。 - 视频文件命名规范:视频存储在 S3、GCS 或本地时应命名为
vid_<youtube_id>.mp4(例如vid_8iXdsvgpwc8.mp4)。 - 摘要真值文件:视频级摘要真值位于
data/inputs/summarization_ground_truth.jsonl。
支持的任务与评估指标
- 任务类型:
- 标准标签分类
- 标签分类与优化工作负载
- 视频摘要生成
- 摘要质量评估(使用 LLM 作为评判员)
- 评估指标:
- 性能指标:
- 视频级标签分类:精确率、召回率、F1 分数
- 视频级摘要生成:基于量规的分数(使用 LLM-as-judge 评估)
- 标签分类与优化工作负载:不适用
- 成本指标:API 调用成本
- 效率指标:延迟/吞吐量
- 性能指标:
基准测试框架与模型
- 评估框架:MediaPerf,一个用于评估多模态基础模型视频理解性能的生产就绪框架。
- 评估模型:涵盖 13 个视觉语言模型,包括 AWS Bedrock (Nova, Pegasus, NVIDIA)、Google Vertex AI (Gemini)、OpenAI (GPT) 以及自托管模型 (Qwen)。完整列表详见 Model Reference Guide。
注意事项
- 原始标签列表中的部分标签(如
funny、effective、exciting)因覆盖范围有限或应用不一致,在分析中被省略。 - 数据集仅包含短视频内容,长视频内容(剧集、电影、体育、新闻)尚未包含。
- 当前基准测试仅评估生成式视觉语言模型,尚未覆盖用于嵌入和搜索/检索工作流的编码器模型。
许可信息
- 源代码许可:Apache License 2.0。
- 人工标注的摘要和标签数据许可:Creative Commons Attribution 4.0 International License (CC-BY 4.0)。
参考文献
[1] Zaeem Hussain, Mingda Zhang, Xiaozhong Zhang, Keren Ye, Christopher Thomas, Zuha Agha, Nathan Ong, Adriana Kovashka. "Automatic Understanding of Image and Video Advertisements." Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 1705-1715. Link




