FineVision
收藏资源简介:
FineVision是一个开源的视觉语言数据集,由Hugging Face提供,旨在训练先进的视觉语言模型。该数据集包含了1730万张图像、2430万个样本、8890万轮对话和95亿个答案标记。它汇集了来自200多个来源的数据,具备多模态和多轮对话的特性,能够支持视觉与语言的结合。每一张图像都配有一个文本标题,这有助于模型理解和生成自然语言。在使用FineVision数据集的10项基准测试中,模型性能平均提升了超过20%。
FineVision is an open-source visual language dataset provided by Hugging Face for training advanced visual language models. It contains 17.3 million images, 24.3 million samples, 88.9 million rounds of dialogue, and 9.5 billion answer annotations. The dataset aggregates data from over 200 sources, featuring multimodal and multi-turn dialogue capabilities, supporting the integration of vision and language. Each image is accompanied by a text title, which helps the model understand and generate natural language. FineVision has helped models achieve an average performance improvement of over 20% in 10 benchmark tests.




