遇见数据集

inesriahi/valor32k-avqa-v2

收藏
Hugging Face2026-05-07 更新2026-05-31 收录
官方服务:

资源简介:

Valor32k-AVQA v2.0 是一个开放式的视听问答数据集和基准,包含28,861个视频和225,487个问题-答案对。每个问题都标注了模态标签(视觉、音频或视听)和六个类别之一:描述、动作、计数、时间、位置和相对位置。

Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (`visual`, `audio`, or `audio-visual`) and one of six categories: `description`, `action`, `count`, `temporal`, `location`, and `relative-position`.

提供机构:
inesriahi
二维码
社区交流群
二维码
科研交流群
商业服务