VideoAVE
收藏资源简介:
VideoAVE是一个多属性视频到文本属性值提取数据集,是第一个公开可用的视频到文本电子商务AVE数据集,涵盖了14个不同的领域和172个独特的属性。为了保证数据质量,我们提出了一种后处理CLIP-MoE过滤系统来移除不匹配的视频-产品对,从而得到一个经过精炼的数据集。我们还在VideoAVE上建立了一个全面的基准,以评估几个最先进的视频视觉语言模型在属性条件值预测和开放属性值对提取任务中的表现。结果表明,视频到文本的AVE仍然是一个具有挑战性的问题,特别是在开放场景中,仍有开发更先进的VLMs的空间。
VideoAVE is a multi-attribute video-to-text attribute value extraction dataset, and the first publicly available video-to-text e-commerce AVE dataset, which covers 14 distinct domains and 172 unique attributes. To ensure data quality, we propose a post-processing CLIP-MoE filtering system to eliminate mismatched video-product pairs, resulting in a refined dataset. We further establish a comprehensive benchmark on VideoAVE to evaluate the performance of several state-of-the-art video visual-language models on two tasks: attribute-conditioned value prediction and open attribute value pair extraction. The results demonstrate that video-to-text AVE remains a challenging problem, especially in open scenarios, where there is still ample room for developing more advanced VLMs.

- 1VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models弗吉尼亚理工大学 · 2025年



