HERM-100K
收藏资源简介:
HERM-100K是由中国科学院计算技术研究所创建的综合性多模态数据集,旨在提升多模态大语言模型(MLLMs)在以人为中心场景下的理解能力。该数据集包含超过100,000条多层次的人类中心注释,涵盖图像级密集描述、实例级注释和属性级注释,以提供全面的视觉信息。数据集的创建过程利用了GPT-4V生成多样化的图像来源注释,并通过预定义模板和GPT-4提示进行多任务预训练和指令微调。HERM-100K主要应用于增强MLLMs在复杂人类中心场景中的视觉理解和任务执行能力,旨在解决现有数据集在人类中心视觉理解中的不足。
HERM-100K is a comprehensive multimodal dataset developed by the Institute of Computing Technology, Chinese Academy of Sciences, aiming to enhance the human-centric scene understanding capabilities of multimodal large language models (MLLMs). It contains over 100,000 multi-level human-centric annotations, covering dense image-level descriptions, instance-level annotations and attribute-level annotations to provide comprehensive visual information. The dataset construction process leverages GPT-4V to generate diverse annotations for various image sources, and conducts multi-task pre-training and instruction fine-tuning via predefined templates and GPT-4 prompts. HERM-100K is primarily applied to strengthen the visual understanding and task execution abilities of MLLMs in complex human-centric scenarios, targeting to address the shortcomings of existing datasets in human-centric visual understanding.

- 1HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding中国科学院计算技术研究所 · 2024年



