Occluded Object Detection Dataset
收藏资源简介:
本研究构建了一个大规模的多模态链式思维数据集,包含超过110k个手持遮挡物体的图像-文本对。数据集基于ObMan数据集,引入了结构化的推理过程,包含描述阶段、自我反思阶段和最终决策阶段,以逐步提高对遮挡物体的识别能力。数据集旨在解决视觉语言模型中遮挡对象理解的问题,适用于多模态任务,如物体识别、场景理解等。
This study develops a large-scale multimodal chain-of-thought dataset containing over 110k image-text pairs of hand-held occluded objects. Based on the ObMan dataset, this dataset incorporates a structured reasoning workflow encompassing the description phase, self-reflection phase and final decision-making phase, to progressively improve the recognition capability of occluded objects. This dataset is designed to address the issue of occluded object understanding in vision-language models, and is applicable to multimodal tasks including object recognition, scene understanding, and other related tasks.

- 1OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance中国科学院上海微系统与信息技术研究所 · 2025年



