HKUSTAudio/AudioX-IFcaps
收藏资源简介:
--- license: cc-by-nc-nd-4.0 task_categories: - text-to-audio size_categories: - 1M<n<10M pretty_name: AudioX-IFcaps --- # [ICLR 2026] AudioX-IFcaps: Instruction-Following Audio Caption Dataset <a href="https://zeyuet.github.io/AudioX/" target="_blank"><img src="https://img.shields.io/badge/🌐%20Project%20Page-blue" alt="Project Page"></a> <a href="https://github.com/ZeyueT/AudioX" target="_blank"><img src="https://img.shields.io/badge/💻%20GitHub-ffffff?logo=github&logoColor=181717" alt="GitHub"></a> <a href="https://arxiv.org/pdf/2503.10522" target="_blank"><img src="https://img.shields.io/badge/📄%20Paper-ICLR%202026-red" alt="Paper"></a> **AudioX-IFcaps** (Instruction-Following) is a large-scale, high-quality multimodal dataset designed for training unified audio and music generation models. The dataset contains over **7 million samples** with fine-grained, structured annotations that enable precise control over audio generation, including sound event categories, counts, temporal ordering, and timestamps. ## 📊 Dataset Statistics - **General Audio**: ~1.3m 10-second video-audio clips - **Music**: ~5.7m 10-second video-music clips - **Total Duration**: ~16k hours of audio content ## 📝 Citation If you use this dataset in your research, please cite: ```bibtex @article{tian2025audiox, title={Audiox: Diffusion transformer for anything-to-audio generation}, author={Tian, Zeyue and Jin, Yizhu and Liu, Zhaoyang and Yuan, Ruibin and Tan, Xu and Chen, Qifeng and Xue, Wei and Guo, Yike}, journal={arXiv preprint arXiv:2503.10522}, year={2025} } @inproceedings{tian2025vidmuse, title={Vidmuse: A simple video-to-music generation framework with long-short-term modeling}, author={Tian, Zeyue and Liu, Zhaoyang and Yuan, Ruibin and Pan, Jiahao and Liu, Qifeng and Tan, Xu and Chen, Qifeng and Xue, Wei and Guo, Yike}, booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference}, pages={18782--18793}, year={2025} } ``` ## 🔗 Related Resources - **Paper**: <a href="https://arxiv.org/pdf/2503.10522" target="_blank">AudioX: Diffusion Transformer for Anything-to-Audio Generation</a> (Accepted to ICLR 2026) - **Project Page**: <a href="https://zeyuet.github.io/AudioX/" target="_blank">https://zeyuet.github.io/AudioX/</a> - **Code**: <a href="https://github.com/ZeyueT/AudioX" target="_blank">GitHub Repository</a> --- **Note**: This dataset is part of the AudioX project. For more information, please refer to the paper and project page.
license: 知识共享署名-非商业性使用-禁止演绎4.0国际许可协议(CC BY-NC-ND 4.0) task_categories: - 文本到音频(text-to-audio) size_categories: - 100万<样本数<1000万 pretty_name: AudioX-IFcaps --- # [ICLR 2026] AudioX-IFcaps:指令遵循式音频字幕数据集 <a href="https://zeyuet.github.io/AudioX/" target="_blank"><img src="https://img.shields.io/badge/🌐%20项目主页-blue" alt="项目主页"></a> <a href="https://github.com/ZeyueT/AudioX" target="_blank"><img src="https://img.shields.io/badge/💻%20GitHub-ffffff?logo=github&logoColor=181717" alt="GitHub仓库"></a> <a href="https://arxiv.org/pdf/2503.10522" target="_blank"><img src="https://img.shields.io/badge/📄%20论文-ICLR%202026-red" alt="论文"></a> **AudioX-IFcaps(指令遵循式音频字幕数据集)**是一款大规模、高质量的多模态数据集,专为训练统一的音频与音乐生成模型而设计。该数据集包含超过700万个样本,配有细粒度、结构化的标注,可实现对音频生成的精准控制,涵盖声音事件类别、数量、时间顺序与时戳等信息。 ## 📊 数据集统计 - **通用音频**:约130万个10秒时长的视频音频片段 - **音乐**:约570万个10秒时长的视频音乐片段 - **总时长**:约1.6万小时的音频内容 ## 📝 引用规范 若您在研究中使用本数据集,请引用如下文献: bibtex @article{tian2025audiox, title={AudioX:面向任意到音频生成的扩散Transformer}, author={Tian, Zeyue and Jin, Yizhu and Liu, Zhaoyang and Yuan, Ruibin and Tan, Xu and Chen, Qifeng and Xue, Wei and Guo, Yike}, journal={arXiv预印本 arXiv:2503.10522}, year={2025} } @inproceedings{tian2025vidmuse, title={VidMuse:一款具备长短时建模能力的简易视频到音乐生成框架}, author={Tian, Zeyue and Liu, Zhaoyang and Yuan, Ruibin and Pan, Jiahao and Liu, Qifeng and Tan, Xu and Chen, Qifeng and Xue, Wei and Guo, Yike}, booktitle={计算机视觉与模式识别会议论文集}, pages={18782--18793}, year={2025} } ## 🔗 相关资源 - **论文**:<a href="https://arxiv.org/pdf/2503.10522" target="_blank">《AudioX:面向任意到音频生成的扩散Transformer》</a>(已被ICLR 2026收录) - **项目主页**:<a href="https://zeyuet.github.io/AudioX/" target="_blank">https://zeyuet.github.io/AudioX/</a> - **代码**:<a href="https://github.com/ZeyueT/AudioX" target="_blank">GitHub仓库</a> --- **注**:本数据集为AudioX项目的组成部分。如需了解更多信息,请参阅相关论文与项目主页。



