VietTravelVQA
收藏资源简介:
Objective: This dataset is designed to benchmark and improve Visual Question Answering (VQA) systems in the context of Vietnamese tourism and cultural heritage. It addresses the lack of high-quality, regionally specific multimodal data for Southeast Asia. Data Content: The dataset comprises thousands of images sourced from Wikimedia Commons, paired with human-verified question-answer sets in Vietnamese. The questions cover five levels of complexity, ranging from basic object identification to deep cultural reasoning. Methodology: 1. Sourcing: Legally compliant images were filtered from Wikimedia Commons. 2. Annotation: Expert annotators generated QA pairs, focusing on architectural details, historical significance, and spatial reasoning. 3. Validation: Data was cleaned using automated scripts to ensure 100% synchronization between metadata (JSON) and image files, with factual auditing via Large Multimodal Models (LMMs). Usage: The data is split into train and test sets (.json). It is intended for training, fine-tuning, and evaluating Vision-Language Models (VLMs) on localized cultural contexts.
任务目标:本数据集旨在针对越南旅游与文化遗产场景下的视觉问答(Visual Question Answering, VQA)系统开展基准测试与性能优化,以解决东南亚地区缺乏高质量、区域专属多模态数据的行业痛点。 数据内容:本数据集包含数千张源自维基共享资源(Wikimedia Commons)的图像,并搭配经人工校验的越南语问答对。问题覆盖五个复杂度层级,从基础的物体识别任务逐步延伸至深层次的文化推理任务。 研究方法: 1. 数据采集:从维基共享资源(Wikimedia Commons)筛选合规合法的图像资源。 2. 标注流程:由专业标注人员生成问答对,重点聚焦建筑细节、历史意义与空间推理维度。 3. 数据验证:通过自动化脚本完成数据清洗工作,确保元数据(JSON格式)与图像文件完全同步,并借助大型多模态模型(Large Multimodal Models, LMMs)开展事实性审核。 使用场景:本数据集已划分为训练集与测试集(均为JSON格式),可用于面向本地化文化场景的视觉语言模型(Vision-Language Models, VLMs)的训练、微调与评估工作。




