AnyCapDataset (ACD)
收藏资源简介:
AnyCap数据集(ACD)是一个大规模的跨模态数据集,涵盖了图像、视频和音频三种模态,包含了28种用户指令和30万个高质量的数据条目。每个样本都包括定制的控制信号和一对字幕,一个是优选的,另一个是被拒绝的,明确反映了控制质量的不同。数据集的构建过程采用了严格的验证和审查阶段,确保了字幕与用户指令的一致性,并且没有幻觉,符合格式要求,并保留了输入的独特特征。该数据集旨在解决可控多模态字幕生成中训练数据不足的问题,并支持学习与各种用户偏好进行精确对齐。
AnyCap Dataset (ACD) is a large-scale cross-modal dataset covering three modalities: image, video, and audio, which contains 28 types of user instructions and 300,000 high-quality data entries. Each sample includes a customized control signal and a pair of captions: one preferred and the other rejected, explicitly reflecting the differences in control quality. The dataset construction process adopts strict verification and review stages to ensure that captions are consistent with user instructions, free of hallucinations, comply with format requirements, and preserve the unique characteristics of the input. This dataset aims to address the shortage of training data for controllable multimodal caption generation, and supports learning to precisely align with various user preferences.
AnyCap数据集概述
数据集基本信息
- 名称:AnyCap
- 类型:多模态可控字幕生成数据集
- 模态支持:图像、音频、视频
- 许可证:MIT License
数据集组成
训练数据集 (AnyCapDataset)
- 状态:即将发布
- 内容:
- 包含图像、音频、视频三种模态数据
- 提供完整的文本标注
- 视频数据需从Hugging Face下载并放置到指定目录
评估基准 (AnyCapEval)
- 状态:已发布
- 获取方式:Hugging Face下载
- 特点:
- 行业级多模态评估基准
- 包含内容评分和风格评分双指标
- 提供全面的评估协议
数据集特点
- 统一框架:支持图像/音频/视频字幕生成
- 可控性:支持通过预定义指令控制字幕风格
- 评估体系:
- 自定义评估指标Key-point Density (KPD)
- 提供与人类判断的高相关性验证
使用说明
数据准备
- 下载AnyCapEval基准数据
- 按照目录结构放置数据文件
评估流程
- 生成字幕文件(content.jsonl和style.jsonl)
- 配置评估脚本路径参数
- 运行评估脚本
相关资源
- 模型权重:Hugging Face获取
- 论文:arXiv论文
引用格式
bibtex @misc{ren2025anycapprojectunifiedframework, title={AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning}, author={Yiming Ren and Zhiqiang Lin and Yu Li and Gao Meng and Weiyun Wang and Junjie Wang and Zicheng Lin and Jifeng Dai and Yujiu Yang and Wenhai Wang and Ruihang Chu}, year={2025}, eprint={2507.12841}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2507.12841}, }

- 1AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning清华大学, 上海人工智能实验室, 复旦大学, 香港中文大学 · 2025年



