amalia-llm/CARAVELA
收藏资源简介:
CARAVELA是一个多模态基准数据集,专门用于评估大型视觉语言模型在葡萄牙文化知识方面的表现。该数据集以欧洲葡萄牙语(pt-PT)为官方语言,包含12,983个图像-问题对,基于3,354个独特的葡萄牙文化实体(如历史古迹、美食、人物、艺术和地点)。每个数据项围绕一个文化实体及其图像构建,涵盖五个文化类别(美食、古迹、人物、艺术、地点)、三个知识维度(时间维度:历史时间线和日期;文化维度:象征意义、社会影响和传统;空间维度:地理背景和物理特征)和三个任务格式(多项选择、视觉问答、推理)。数据集结构包括图像、问题、选项(对于多项选择任务)、答案(正确选项字母或自由文本)、实体名称、类别标签、知识维度区域和支持文本(来自葡萄牙维基百科)。注释部分使用Apache-2.0许可,图像部分源自Wikimedia Commons并保留原始许可。该数据集旨在量化模型在葡萄牙文化元素上的对齐和视觉理解能力,提供细粒度的诊断评估。
CARAVELA is a multimodal benchmark for evaluating the Portuguese cultural knowledge of large vision-language models (LVLMs). The official benchmark language is exclusively European Portuguese (pt-PT). It contains 12,983 image–question pairs derived from 3,354 unique Portuguese cultural entities. Each item is anchored to a culturally relevant Portuguese entity and its image, organized along three axes: five cultural categories (Gastronomy, Monuments, Personalities, Art, Locations), three knowledge dimensions (Temporal: historical timelines and dates; Cultural: symbolic meaning, societal impact, and traditions; Spatial: geographical context and physical features), and three task formats (Multiple-Choice, Visual Question Answering, and Reasoning). The dataset structure includes fields such as image, question, options (for MCQ), answer (correct option letter or free-text), entity, category, area, and supporting text (from Portuguese Wikipedia). Annotations are licensed under Apache-2.0, while images are sourced from Wikimedia Commons and retain their original licenses. It provides a rigorous framework to quantify cultural alignment and visual understanding across Portugals cultural elements.




