lance-format/textvqa-lance
收藏资源简介:
TextVQA(Lance格式)是TextVQA数据集的Lance格式版本,主要用于视觉问答任务,要求模型能够阅读图像中的文本。数据集包含图像字节、问题、10个参考答案、OCR识别的文本标记以及CLIP图像和问题嵌入。数据集分为训练集(34,602行)和验证集(5,000行)。提供了详细的列描述和预构建的索引,支持跨模态文本→图像搜索。数据集采用CC BY 4.0许可证,由Singh等人(Facebook AI Research)发布。
TextVQA (Lance Format) is a Lance-formatted version of the TextVQA dataset, designed for visual question answering tasks that require reading text in the image. The dataset includes image bytes, questions, 10 reference answers, OCR tokens detected on the image, and CLIP image and question embeddings. It is divided into training (34,602 rows) and validation (5,000 rows) splits. The dataset provides detailed column descriptions and pre-built indices, supporting cross-modal text→image search. It is released under CC BY 4.0 by Singh et al. (Facebook AI Research).




