textvqa
收藏官方服务:
资源简介:
TextVQA数据集专注于提升模型在图像中阅读和理解文本的能力,以回答相关问题。它包含从OpenImages数据集中筛选出的28408张图像,并针对这些图像提出了45336个问题。每个问题都配有10个人工标注的答案,并采用VQA准确率作为评估指标。数据集提供图像、问题、答案以及其他元数据,支持视觉问答任务。该数据集使用CC-BY 4.0协议授权。
The TextVQA dataset focuses on enhancing models' ability to read and understand text within images to answer relevant questions. It contains 28,408 images filtered from the OpenImages dataset, with 45,336 questions formulated for these images. Each question is paired with 10 manually annotated answers, and VQA accuracy is adopted as the evaluation metric. The dataset provides images, questions, answers and other metadata, supporting visual question answering tasks. This dataset is licensed under the CC-BY 4.0 license.
提供机构:
facebook创建时间:
2024-07-19
搜集汇总
数据集介绍

背景与挑战
背景概述
TextVQA数据集专注于提升模型在图像中阅读和理解文本的能力,以回答相关问题。它包含28,408张来自OpenImages的图像和45,336个问题,每个问题提供10个人工标注答案,并使用VQA准确率作为评估指标,采用CC-BY 4.0协议授权。
以上内容由遇见数据集搜集并总结生成



