CultureVerse
收藏资源简介:
CultureVerse是一个大规模的多模态文化理解基准数据集,由澳门大学、微软研究院等机构联合构建。该数据集涵盖了19,682个文化概念,覆盖188个国家和地区,包含15个文化类别和3种问题类型,总样本量达到228,053条。数据集的构建过程结合了自动化网络爬取和专家人工标注,确保了数据的多样性和高质量。CultureVerse旨在评估和提升视觉-语言模型(VLMs)的多文化理解能力,特别关注全球南方国家和少数民族文化的代表性。通过该数据集,研究者可以训练和评估模型在不同文化背景下的表现,推动更公平、更具文化意识的多模态AI系统的发展。
CultureVerse is a large-scale multimodal cultural understanding benchmark dataset jointly constructed by institutions including the University of Macau and Microsoft Research. This dataset covers 19,682 cultural concepts, spans 188 countries and regions, includes 15 cultural categories and 3 question types, with a total sample size of 228,053. The dataset's construction combines automated web crawling and expert manual annotation, ensuring the diversity and high quality of the data. CultureVerse aims to evaluate and enhance the multicultural understanding capabilities of Vision-Language Models (VLMs), with a particular focus on the representation of Global South countries and ethnic minority cultures. Through this dataset, researchers can train and evaluate model performance across different cultural contexts, promoting the development of more equitable and culturally aware multimodal AI systems.




