weamontt/VQA_Food_Dataset
收藏资源简介:
该数据集是一个越南食品视觉问答(VQA)数据集,设计用于在越南食品图像上进行视觉问答任务。它由图像-问题-答案三元组组成,每个问题以越南语书写,需要理解视觉内容和语义信息。数据集支持训练和评估结合计算机视觉和自然语言处理的多模态模型。数据集包含6,034个样本,没有缺失图像,问题平均长度为9.57个词。答案分布相对平衡,最常见的答案占比31%。问题类型包括是/否、识别和其他类型。数据集结构包含图像和注释文件,数据格式为JSON。数据增强技术包括图像和文本增强,以提高泛化能力。数据集质量良好,但存在少量重复QA对和是/否问题占主导的局限性。
This dataset is designed for the Visual Question Answering (VQA) task on Vietnamese food images. It consists of image–question–answer triplets, where each question is written in Vietnamese and requires understanding both visual content and semantic information. The dataset supports training and evaluation of multimodal models combining computer vision and natural language processing. It contains 6,034 samples with no missing images, and the average question length is 9.57 words. The answer distribution is relatively balanced, with the most frequent answer accounting for 31% of the dataset. Question types include Yes/No, Recognition, and Other. The dataset structure includes images and annotation files, with data stored in JSON format. Data augmentation techniques are applied to both images and text to improve generalization. The dataset quality is good, but there are some limitations such as a small number of duplicate QA pairs and a dominance of Yes/No questions.




