Phantom-Data
收藏资源简介:
Phantom-Data是一个通用主题一致的视视频生成数据集,旨在解决现有模型在忠实遵循文本指令方面的挑战。该数据集包含约一百万个身份一致的跨类别对,通过一个三阶段的流程构建:一个通用的输入对齐的主题检测模块,从超过5300万个视频和30亿张图像中进行大规模跨上下文主题检索,以及基于先验的视觉一致性验证。该数据集显著提高了提示对齐和视觉质量,同时保持了与成对基线相当的身份一致性。
Phantom-Data is a general thematic consistent video and image generation dataset designed to address the challenges existing models face in adhering to textual instructions accurately. The dataset comprises approximately one million consistent identity pairs across categories, constructed through a three-phase process: a general input alignment thematic detection module, large-scale cross-contextual thematic retrieval from over 53 million videos and 3 billion images, and a prior-based visual consistency validation. This dataset significantly enhances prompt alignment and visual quality while maintaining a comparable identity consistency with paired baselines.
Phantom-Data 数据集概述
基本信息
- 数据集名称: Phantom-Data
- 开发团队: Intelligent Creation Lab, ByteDance
- 主要贡献者: Zhuowei Chen, Bingchuan Li, Tianxiang Ma, Lijie Liu 等
- 论文状态: arXiv preprint (2025)
- 论文编号:
- Phantom-Data: arXiv:2506.18851
- Phantom: arXiv:2502.11079
数据集目标
- 解决主题到视频生成中的"复制粘贴"问题
- 提供首个通用的大规模跨对主题一致视频生成数据集
核心特点
- 规模: 包含约100万个身份一致的对
- 多样性: 覆盖30,000多个多主题场景
- 质量: 高视觉一致性
数据集构建流程
- 主题中心检测模块: 获取通用且输入对齐的主题
- 跨上下文检索系统: 从5300万视频和30亿图像中检索跨对候选
- 先验引导身份验证: 确保上下文变化下的视觉一致性
研究成果
- 相比基线方法显著改善了提示对齐和视觉质量
- 在保持身份一致性方面与基线方法相当
非合成数据选择原因
- 现有SOTA模型(GPT4o和DreamO)仍会产生不一致主题
- 跨对数据构建流程能提供完全相同的主题
引用信息
bibtex @article{chen2025phantom-data, title={Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset}, author={Chen, Zhuowei and Li, Bingchuan and Ma, Tianxiang and Liu, Lijie and Liu, Mingcong and Zhang, Yi and Li, Gen and Li, Xinghui and Zhou, Siyu and He, Qian and Wu, Xinglong}, journal={arXiv preprint arXiv:2506.18851}, year={2025} } @article{liu2025phantom, title={Phantom: Subject-consistent video generation via cross-modal alignment}, author={Liu, Lijie and Ma, Tianxiang and Li, Bingchuan and Chen, Zhuowei and Liu, Jiawei and Li, Gen and Zhou, Siyu and He, Qian and Wu, Xinglong}, journal={arXiv preprint arXiv:2502.11079}, year={2025} }
许可信息
- 许可证: Creative Commons Attribution-ShareAlike 4.0 International License

- 1Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset字节跳动智能创作实验室 · 2025年



