VisualQuest
收藏资源简介:
VisualQuest是一个由大连理工大学计算机科学与技术学院和新疆师范大学计算机科学与技术学院合作创建的图像数据集。该数据集包含3529张分为四个主题类别的图像,这些类别包括公众人物、流行媒体、语言表达和文学作品。每类图像都采用艺术性的表现手法,如表情符号、插图、漫画等,增加了识别任务的复杂性。数据集经过精心策划,旨在评估LLM在识别融入了领域特定知识和抽象视觉推理的非传统图像方面的能力,为多模态推理和模型架构设计的研究提供了宝贵的基准。
VisualQuest is an image dataset co-created by the School of Computer Science and Technology, Dalian University of Technology and the School of Computer Science and Technology, Xinjiang Normal University. This dataset contains 3529 images categorized into four thematic classes: public figures, popular media, linguistic expressions, and literary works. Each class features images created using artistic expression techniques including emojis, illustrations, comics and other styles, which increases the complexity of the recognition task. The dataset is meticulously curated to evaluate the ability of Large Language Models (LLMs) to recognize non-traditional images that incorporate domain-specific knowledge and abstract visual reasoning, providing a valuable benchmark for research on multimodal reasoning and model architecture design.




