BabyVLM-V2
收藏资源简介:
BabyVLM-V2是由波士顿大学和索尼集团公司联合开发的婴儿启发的视觉语言建模框架,旨在通过发展心理学原理进行样本高效的预训练。该数据集包含768,000条图像-话语对,以及181,000条视频-话语对和63,000条交错序列,数据来源于SAYCam的婴儿中心视角的纵向视听语料库。数据集创建过程最大限度地减少了人工干预,以保持儿童感官摄入的真实性。BabyVLM-V2的应用领域主要集中在发展合理的视觉基础模型预训练,旨在解决早期儿童感知能力的模拟和评估问题。
BabyVLM-V2 is a baby-inspired visual-language modeling framework co-developed by Boston University and Sony Group Corporation, which aims to conduct sample-efficient pre-training based on developmental psychology principles. This dataset contains 768,000 image-utterance pairs, 181,000 video-utterance pairs and 63,000 interleaved sequences, sourced from the longitudinal audio-visual corpus with infant-centered perspectives from SAYCam. The dataset creation process minimizes human intervention to preserve the authenticity of children's sensory inputs. The main application scenarios of BabyVLM-V2 focus on pre-training for developing valid visual foundation models, aiming to solve the problems of simulating and evaluating early childhood perceptual abilities.




