LLaVAR-2
收藏资源简介:
LLaVAR-2是一个用于增强多模态对齐的高质量文本丰富图像指令调优数据集,由布法罗大学和Adobe研究院创建。该数据集包含42万条详细丰富的描述性字幕和38.2万条视觉问答数据对,通过GPT-4o自动生成。数据集的创建过程结合了人工注释和大型语言模型的混合指令生成,确保了数据的高质量和多样性。LLaVAR-2主要用于提升多模态大语言模型在处理文本丰富图像任务中的表现,旨在解决现有数据集在文本理解能力上的不足。
LLaVAR-2 is a high-quality text-rich image instruction-tuning dataset for advancing multimodal alignment, developed by the University at Buffalo and Adobe Research. This dataset includes 420,000 detailed and rich descriptive captions and 382,000 visual question-answering (VQA) pairs, which are automatically generated via GPT-4o. The dataset's construction process integrates hybrid instruction generation that combines human annotations and large language models, ensuring the high quality and diversity of the collected data. LLaVAR-2 is primarily utilized to improve the performance of multimodal large language models in tasks involving text-rich images, with the goal of addressing the deficiencies of existing datasets in terms of text understanding capabilities.

- 1A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation布法罗大学, Adobe研究院 · 2024年



