DenseWorld-1M
收藏资源简介:
DenseWorld-1M是第一个大规模、详细、密集的接地字幕数据集,旨在填补现有数据集在详细描述、关系和大量对象描述方面的不足。它通过三阶段标注流程生成,包括开放世界感知、详细对象字幕生成和密集字幕合并。
DenseWorld-1M is the first large-scale, detailed, dense grounded caption dataset. It is designed to address the shortcomings of existing datasets in terms of detailed descriptions, relational semantics and descriptions of numerous objects. The dataset is created through a three-stage annotation pipeline, which includes open-world perception, detailed object caption generation and dense caption merging.
DenseWorld-1M 数据集概述
基本描述
- 数据集名称: DenseWorld-1M
- 目标: 提供首个大规模、详细、密集的接地字幕数据集,用于现实世界场景理解。
- 特点: 包含视觉实体的地面位置和关系,提供详细描述和大量对象描述。
数据集构建
- 标注流程: 三阶段标注管道
- 开放世界感知: 获取实体级掩码和标签。
- 详细对象字幕生成: 在第一阶段掩码和标签的指导下生成对象级详细字幕。
- 密集字幕合并: 将对象字幕和掩码合并为空间和关系密集字幕。
- 辅助模型:
- 详细区域字幕模型 (Detailed Region Caption model)
- 空间字幕合并模型 (Spatial Caption Merging model)
技术细节
- 应用场景:
- 视觉语言理解
- 视觉接地
- 区域字幕生成
- 图像分辨率: 高分辨率图像
当前状态
- 开放进度:
- 数据集清理中
- 开源流程正在审核
- 计划2024年7月底前在HuggingFace上完整开放
- 待完成事项:
- 发布不同模型的训练代码
- 发布数据集
相关资源
- 论文: arXiv:2506.24102
- 代码库: GitHub仓库
- 数据集托管: HuggingFace
引用信息
bibtex @misc{li2025denseworld1m, title={DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World}, author={Xiangtai Li and Tao Zhang and Yanwei Li and Haobo Yuan and Shihao Chen and Yikang Zhou and Jiahao Meng and Yueyi Sun and Shilin Xu and Lu Qi and Tianheng Cheng and Yi Lin and Zilong Huang and Wenhao Huang and Jiashi Feng and Guang Shi}, year={2025}, eprint={2506.24102}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2506.24102}, }




