AudioMapCap-44K
收藏资源简介:
AudioMapCap-44K是首个面向时间感知密集音频描述任务的数据集,由香港科技大学(广州)与快手科技联合构建。该数据集包含43,870条音频-描述对,总时长约769.7小时,每条描述均涵盖精细的声学属性及精确的时间边界。数据集通过精心设计的标注流程生成,旨在为模型提供多事件、多属性、多关系的密集描述监督信号。其应用领域聚焦于提升音频理解模型在细粒度属性覆盖与事件-时间对齐方面的能力,以解决传统方法无法同时兼顾细粒度与时间感知的难题。
AudioMapCap-44K is the first dataset dedicated to the task of time-aware dense audio captioning, jointly constructed by The Hong Kong University of Science and Technology (Guangzhou) and Kuaishou Technology. This dataset consists of 43,870 audio-description pairs with a total duration of approximately 769.7 hours, where each description covers fine-grained acoustic attributes and precise temporal boundaries. Generated via a meticulously designed annotation pipeline, the dataset aims to provide audio models with dense supervisory signals that support multi-event, multi-attribute and multi-relation scenarios. Its targeted application scenarios focus on enhancing the capability of audio understanding models in terms of fine-grained attribute coverage and event-time alignment, so as to solve the long-standing challenge that traditional audio understanding approaches fail to simultaneously balance both fine-grained detail and temporal awareness.
数据集概述:AudioMapCap-44K
AudioMapCap-44K 是一个专为时间感知密集音频字幕生成(Time-Aware Dense Audio Captioning, TDAC)任务构建的首个细粒度音频字幕数据集。该数据集由 AudioMap 框架配套推出,旨在解决现有音频字幕方法在细粒度属性描述和时间边界对齐方面的不足。
核心特点
- 规模:包含 44,000 条精心标注的字幕,覆盖多事件、多属性、多关系的音频描述。
- 时间感知:每条字幕不仅描述音频内容的语义,还提供精确的时间边界信息,支持对音频事件的时间定位。
- 密集描述:字幕包含丰富的细粒度属性(如声学维度、事件类型、关系等),而非简单的整体概括。
- 训练支持:专为强化学习(RL)框架设计,支持基于“完形填空与选择”(cloze-and-choice)奖励范式的训练。
用途与任务
该数据集服务于 TDAC 任务,即要求模型在生成多事件、多属性、多关系的详细音频描述时,同步输出时间戳以定位事件。它可用于训练和评估音频字幕生成模型,尤其适合需要精细时间监督和细粒度语义理解的研究场景。
获取方式
数据集目前标注为“即将发布”(Coming Soon),计划托管于 Hugging Face 平台(🤗 Dataset 链接待上线)。同时,相关项目主页已开放(https://ryysayhi.github.io/AudioMap/ ),论文预印本也将于近期发布(arXiv 状态:Coming Soon)。





