PixLore
收藏资源简介:
PixLore数据集是由高知特瓦伦西亚的研究人员Diego Bonilla Salvador创建,包含100,000张来自COCO数据集的图像。该数据集通过结合多种计算机视觉模型和ChatGPT的增强,生成了详细且丰富的图像描述。创建过程中,每张图像都经过多个先进的计算机视觉模型的处理,最终通过ChatGPT生成文本描述。PixLore数据集主要用于图像描述任务,旨在通过小规模模型实现复杂的图像理解,解决现有模型在描述细节和上下文方面的不足。
The PixLore dataset was created by researcher Diego Bonilla Salvador from Cognizant Valencia, and includes 100,000 images sourced from the COCO dataset. It generates detailed and rich image captions by integrating multiple computer vision models with ChatGPT-powered enhancements. During its development, each image was processed by several state-of-the-art computer vision models, before final textual descriptions were generated via ChatGPT. Primarily intended for image captioning tasks, the PixLore dataset aims to enable complex image understanding using small-scale models, and addresses the limitations of existing models in describing image details and contextual information.




