OmniVQA
收藏资源简介:
OmniVQA数据集是首个开源的全向视觉问答数据集,基于ERP格式的全景图像构建,包含三种任务类型:物体识别、属性分析和空间关系推理,尤其关注极地区域。该数据集由香港理工大学、香港科技大学等研究机构合作开发,旨在为全向视觉问答提供全面的数据支持和基准测试。数据集包含1213张全景图像,总计4852个问答对,通过迭代优化策略生成,以确保数据质量和模型性能。OmniVQA数据集及其配套的基准测试OmniVQABench,为评估和改进全向视觉问答模型提供了重要的工具。
The OmniVQA dataset is the first open-source omnidirectional visual question answering (VQA) dataset constructed based on equirectangular projection (ERP) format panoramic images. It encompasses three task types: object recognition, attribute analysis, and spatial relationship reasoning, with a particular focus on polar regions. This dataset was collaboratively developed by research institutions including The Hong Kong Polytechnic University and The Hong Kong University of Science and Technology, aiming to provide comprehensive data support and benchmarking for omnidirectional VQA. The dataset contains 1,213 panoramic images and a total of 4,852 question-answer pairs, which are generated via an iterative optimization strategy to ensure data quality and model performance. The OmniVQA dataset and its supporting benchmark, OmniVQABench, serve as critical tools for evaluating and improving omnidirectional VQA models.




