CoSyn400K
收藏资源简介:
CoSyn400K是一个大规模的合成视觉语言数据集,由宾夕法尼亚大学和艾伦人工智能研究所的研究人员创建。该数据集包含400K个图像和270万行视觉语言指令调优数据,旨在帮助视觉语言模型更好地理解和处理富含文本的图像。数据集通过利用编码能力强大的文本大型语言模型自动生成合成数据,涵盖了图表、文档、数学问题、表格、图表、向量图形、乐谱、电路图和化学结构等九类文本丰富的图像。
CoSyn400K is a large-scale synthetic vision-and-language dataset created by researchers from the University of Pennsylvania and the Allen Institute for AI. This dataset contains 400K images and 2.7 million lines of vision-and-language instruction tuning data, aiming to help vision-and-language models better understand and process text-rich images. The dataset automatically generates synthetic data by leveraging large language models (LLMs) with strong coding capabilities for text, covering nine categories of text-rich images including charts, documents, mathematical problems, tables, charts, vector graphics, sheet music, circuit diagrams, and chemical structures.

- 1Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation宾夕法尼亚大学,艾伦人工智能研究所 · 2025年



