VisualWebInstruct
收藏资源简介:
VisualWebInstruct是一个由滑铁卢大学等机构提出的新型数据集,通过利用搜索引擎创建包含多个学科如数学、物理、金融、化学等的高质量、多样化的数据集。该数据集从30,000个精选的种子图像出发,使用Google Image搜索来识别包含相似图像的网站,收集并处理超过700,000个独立URL源的HTML内容,构建了一个大约有900,000个问答对的数据集,其中40%是视觉问答对,其余为文本问答对。该数据集适用于提升视觉语言模型在需要多步骤推理的复杂任务上的性能。
VisualWebInstruct is a novel dataset proposed by institutions including the University of Waterloo. It is developed to build high-quality and diverse datasets spanning multiple disciplines such as mathematics, physics, finance, chemistry and more via search engines. Starting from 30,000 carefully selected seed images, this dataset uses Google Image Search to identify websites containing similar images, collects and processes HTML content from over 700,000 independent URL sources, and constructs a dataset with approximately 900,000 question-answer pairs, of which 40% are visual question-answer pairs and the remaining are text-based question-answer pairs. This dataset is designed to enhance the performance of vision-language models on complex tasks that require multi-step reasoning.

- 1VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search滑铁卢大学, 多伦多大学, 圣塔巴巴拉加州大学, 卡内基梅隆大学, 新加坡国立大学, 独立研究者, Netmind.ai · 2025年



