AudioChaps-Alignment; AudioChaps-CoT; AudioChaps-Eval
收藏资源简介:
AudioChaps数据集集合由萨里大学与华为诺亚方舟实验室联合创建,是首个专为音频章节化任务设计的综合性数据集,旨在通过强化学习对齐大型音频语言模型与人类编辑判断。该集合包含三个子集:AudioChaps-Alignment来源于YouTube创作者标注的章节边界,覆盖结构化语音、动态媒体、游戏及音乐四种声学场景;AudioChaps-CoT通过音频到文本模态桥接管道生成结构化推理轨迹,提供基于证据的监督信号;AudioChaps-Eval作为首个纯音频章节化基准测试集,用于评估模型泛化能力。数据集创建过程分为三个阶段:边界相关伪思维链生成、声学感知日志净化及最终合成,确保推理格式规范且证据充分。该数据集旨在解决媒体内容章节化这一关键挑战,推动音频流从连续无序向结构可导航的转型,为媒体编目、索引及分发等实际应用奠定基础。
The AudioChaps dataset collection, co-developed by the University of Surrey and Huawei Noah's Ark Lab, is the first comprehensive dataset specifically designed for audio chapterization tasks, aiming to align large audio language models with human editorial judgments via reinforcement learning. This collection includes three subsets: AudioChaps-Alignment, derived from chapter boundaries annotated by YouTube creators, covers four acoustic scenarios: structured speech, dynamic media, gaming, and music; AudioChaps-CoT generates structured reasoning trajectories through an audio-to-text modality bridging pipeline, providing evidence-based supervision signals; AudioChaps-Eval, as the first pure audio chapterization benchmark dataset, is used to evaluate model generalization capabilities. The dataset creation process consists of three stages: boundary-related pseudo-Chain-of-Thought generation, acoustic-aware log purification, and final synthesis, ensuring standardized reasoning formats and sufficient evidence. This dataset aims to address the critical challenge of media content chapterization, promoting the transformation of audio streams from continuous and unstructured to structurally navigable, and laying a foundation for practical applications such as media cataloging, indexing, and distribution.
AudioChaps 数据集详情
名称:AudioChaps
类型:音频大语言模型的后训练框架
核心功能:该框架基于 GRPO(Group Relative Policy Optimization)算法,用于将现有的大型音频语言模型与创作者编写的编辑判断进行对齐,以支持媒体章节化(media chapterization)任务。
特点:
- 采用强化学习/策略优化方法(GRPO)进行模型微调与行为对齐。
- 聚焦于媒体内容的自动章节切分与标注,服务于视频/音频的内容组织与导航。
用途:用于改进音频语言模型在媒体章节化任务上的表现,使其输出更符合创作者的编辑意图。

- 1Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization萨里大学; 华为诺亚方舟实验室 · 2026年



