SMIR
收藏资源简介:
SMIR数据集由加州大学伯克利分校、斯坦福大学和加州理工学院的研究团队开发,旨在解决多图像推理任务中的数据集稀缺问题。该数据集包含160,000个训练样本,通过多模态嵌入技术提取高度相关的图像,并结合开源大语言模型生成高质量的指令数据。数据集生成过程包括图像和文本的多模态嵌入构建、聚类算法以及基于开源模型的指令生成。SMIR数据集的应用领域主要集中在多图像推理任务中,旨在提升视觉-语言模型在多图像场景下的推理能力,解决现有开源模型在多图像任务中表现不佳的问题。
The SMIR dataset was developed by research teams from the University of California, Berkeley, Stanford University, and the California Institute of Technology, with the goal of addressing the scarcity of datasets for multi-image reasoning tasks. It contains 160,000 training samples, where highly correlated images are extracted using multimodal embedding technologies, and high-quality instruction data is generated in combination with open-source large language models. The dataset generation process involves the construction of multimodal embeddings for both images and text, clustering algorithms, and instruction generation based on open-source models. The SMIR dataset is primarily applied to multi-image reasoning tasks, aiming to improve the reasoning abilities of vision-language models in multi-image scenarios and resolve the underperformance issue of existing open-source models on multi-image tasks.




