InsAVE-80K
收藏资源简介:
InsAVE-80K是由北京智源人工智能研究院与北京大学联合构建的首个大规模指令引导音视频联合编辑数据集,旨在为开放世界音视频内容协同操控提供高质量数据基础。该数据集包含约8万条样本,涵盖79K训练对和1K评估对,每条数据均包含源媒体、合成目标及文本指令三元组,数据源自公开在线平台及多个权威音视频数据集,并经过多阶段严格筛选。其创建过程采用可扩展数据合成流水线,通过掩码引导编辑引擎自动生成指令与目标,并辅以多模态大模型评估与人工验证确保数据可靠性。该数据集主要应用于音视频联合生成与编辑领域,致力于解决指令引导下跨模态细粒度内容同步修改的难题,推动可控媒体内容创作的技术发展。
InsAVE-80K is the first large-scale instruction-guided audio-visual joint editing dataset jointly constructed by the Beijing Academy of Artificial Intelligence (BAAI) and Peking University. It aims to provide a high-quality data foundation for collaborative manipulation of open-world audio-visual content. This dataset contains approximately 80,000 samples, including 79K training pairs and 1K evaluation pairs. Each data entry includes a triplet of source media, synthesized target, and text instruction. The data is sourced from public online platforms and multiple authoritative audio-visual datasets, and has undergone multi-stage strict screening. Its creation process adopts a scalable data synthesis pipeline, which automatically generates instructions and targets via a mask-guided editing engine, and is supplemented by multi-modal large model evaluations and manual verification to ensure data reliability. This dataset is mainly applied in the field of audio-visual joint generation and editing, and is committed to solving the challenge of synchronized modification of cross-modal fine-grained content under instruction guidance, so as to promote the technological development of controllable media content creation.




