遇见数据集

AVUTBenchmark

收藏
魔搭社区2026-07-15 更新2026-07-15 收录
官方服务:

资源简介:

# Audio-centric Video Understanding Benchmark (AVUT) This dataset is presented in the paper [Audio-centric Video Understanding Benchmark without Text Shortcut](https://huggingface.co/papers/2503.19951). **Code Repository:** [https://github.com/lark-png/AVUT](https://github.com/lark-png/AVUT) **Paper:** [https://arxiv.org/pdf/2503.19951](https://arxiv.org/pdf/2503.19951) ## Introduction The Audio-centric Video Understanding Benchmark (AVUT) aims to evaluate the video comprehension capabilities of multimodal Large Language Models (LLMs), with a particular focus on auditory information. Audio offers critical context, emotional cues, and semantic meaning that visual data alone often lacks, and AVUT is designed to thoroughly test this aspect. AVUT introduces a suite of carefully designed audio-centric tasks, holistically testing the understanding of both audio content and audio-visual interactions in videos. A key contribution of this benchmark is its approach to the "text shortcut problem," which exists in many other benchmarks where correct answers can be inferred from question text alone without requiring actual video analysis. AVUT addresses this by proposing an answer permutation-based filtering mechanism. ## Dataset Structure The AVUT dataset includes video annotation JSON files essential for evaluation. Specifically, there are two primary annotation files: * `AV_Human_data.json`: Contains annotations meticulously created by human annotators. * `AV_Gemini_data.json`: Contains annotations automatically generated by the Gemini model. These files provide the basis for evaluating and understanding the performance of multimodal LLMs in audio-centric video comprehension tasks.

提供机构:
maas
创建时间:
2026-02-07
二维码
社区交流群
二维码
科研交流群
商业服务