遇见数据集

YqjMartin/FLARE

收藏
Hugging Face2026-05-10 更新2026-05-31 收录
官方服务:

资源简介:

FLARE是一个全模态长视频音视频检索基准测试,包含用户模拟查询。它从Video-MME中筛选了399个长视频(10-60分钟,总计225.4小时),并将其分割成87,697个细粒度片段。每个片段标注了三种类型的描述:仅视觉描述、仅音频描述和统一音视频描述。数据集还提供了274,933个用户模拟查询,包括86,350个仅视觉查询、135,003个仅音频查询和53,580个跨模态查询。评估涵盖两个轴:模态范围(视觉、音频、视觉+音频)和查询机制(基于描述、基于查询),以及四个方向(文本↔片段、文本↔视频)。这是首个在同一个长视频库上同时探索音视频融合和现实用户风格查询的长视频检索基准测试。

FLARE is a Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries. It screens 399 long-form videos (10–60 minutes, 225.4 hours total) from Video-MME and segments them into 87,697 fine-grained clips. Each clip is annotated with three captions: vision-only, audio-only, and unified audiovisual. The dataset includes 274,933 user-simulated queries: 86,350 vision-only queries, 135,003 audio-only queries, and 53,580 cross-modal queries. Evaluation spans two axes: modality scope (vision, audio, vision+audio) and query regime (caption-based, query-based), across four directions (text↔clip, text↔video). To the best of our knowledge, FLARE is the first long-video retrieval benchmark that jointly probes audiovisual fusion and realistic user-style queries on the same long-video gallery.

提供机构:
YqjMartin
二维码
社区交流群
二维码
科研交流群
商业服务