遇见数据集

Low Data Visual Question Answering

收藏
Monash University Figshare2026-02-11 更新2026-07-07 收录
官方服务:

资源简介:

Visual question answering (VQA) is the problem of understanding rich image contexts and answering complex natural language questions about them. VQA models have recently achieved remarkable results when training on large-scale labeled datasets. However, annotating large amounts of data is not feasible in many domains. In this thesis, we address the problem of VQA in low labeled data regimes, which is under-explored in the literature. We leverage natural language's inherent compositional properties to break down the complex questions and learn the sub-questions which are easy to understand. We propose four different approaches to learn sub-questions and provide a strong foundation for learning to answer complex questions with low data. Our results demonstrate significant improvements over the baselines.

创建时间:
2022-07-27
二维码
社区交流群
二维码
科研交流群
商业服务