遇见数据集

AnonymousFLARE/FLARE

收藏
Hugging Face2026-05-07 更新2026-05-31 收录
官方服务:

资源简介:

FLARE是一个全模态长视频视听检索基准数据集,基于用户模拟查询构建。它从Video-MME中筛选出399个长视频(每个视频时长10-60分钟,总时长225.4小时),并将其分割成87,697个细粒度片段。每个片段都标注了三种类型的字幕:仅视觉字幕、仅音频字幕和统一视听字幕。此外,数据集包含274,933个用户模拟查询,分为86,350个仅视觉查询(从视觉字幕重写并经过检索验证)、135,003个仅音频查询(从音频字幕重写并验证)和53,580个跨模态查询(从统一字幕重写,并通过硬双模态约束过滤,确保查询需要视听融合才能识别目标片段)。评估覆盖两个轴:模态范围(视觉、音频、视觉+音频)和查询制度(基于字幕、基于查询),涉及四个方向(文本↔片段、文本↔视频)。该基准旨在联合探究长视频库上的视听融合和现实用户风格查询,是首个针对长视频检索的此类基准。

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries. It screens 399 long-form videos (10-60 minutes each, 225.4 hours total) from Video-MME and segments them into 87,697 fine-grained clips. Each clip is annotated with three captions: vision-only, audio-only, and unified audiovisual. The dataset includes 274,933 user-simulated queries: 86,350 vision-only queries rewritten from vision captions and validated by rank-1 retrieval against the vision gallery, 135,003 audio-only queries rewritten from audio captions and validated similarly, and 53,580 cross-modal queries rewritten from unified captions, filtered by a hard bimodal constraint that ensures only joint vision+audio queries uniquely identify the target clip, isolating evidence requiring audiovisual fusion. Evaluation spans two axes: modality scope (vision, audio, vision+audio) and query regime (caption-based, query-based) across four directions (text↔clip, text↔video). It is the first long-video retrieval benchmark to jointly probe audiovisual fusion and realistic user-style queries on the same long-video gallery.

提供机构:
AnonymousFLARE
二维码
社区交流群
二维码
科研交流群
商业服务