遇见数据集

PDF-MVQA

收藏
arXiv2024-04-19 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

PDF-MVQA数据集是由墨尔本大学、悉尼大学和西澳大学的研究团队开发的,专门针对研究期刊文章的多页和多模态信息检索。该数据集不同于传统的机器阅读理解任务,其主要目标是检索包含答案或视觉丰富文档实体(如表格和图表)的完整段落。数据集包含3146篇文档,总计30,239页,每篇文档平均关联84个问题,共计262,928个问题-答案对。PDF-MVQA数据集通过引入新的多模态文档实体检索框架,旨在提高现有视觉和语言模型在处理文本主导文档在视觉问答任务中的挑战。

The PDF-MVQA dataset is developed by research teams from the University of Melbourne, the University of Sydney, and the University of Western Australia, specifically designed for multi-page and multimodal information retrieval over research journal articles. Unlike traditional machine reading comprehension tasks, its core objective is to retrieve full-length paragraphs that contain either answers or visually-rich document entities such as tables and figures. The dataset consists of 3,146 documents, totaling 30,239 pages, with an average of 84 questions associated with each document, resulting in a total of 262,928 question-answer pairs. By proposing a novel multimodal document entity retrieval framework, the PDF-MVQA dataset aims to enhance the performance of existing visual-language models in addressing the challenges encountered when processing text-dominant documents in visual question answering tasks.

创建时间:
2024-04-19
搜集汇总
数据集介绍
PDF-MVQA 数据集图片
背景与挑战
背景概述
PDF-MVQA数据集是一个专门用于研究期刊文章的多页和多模态信息检索的数据集,由墨尔本大学、悉尼大学和西澳大学的研究团队开发。它包含3146篇文档、30,239页和262,928个问题-答案对,主要目标是检索包含答案或视觉实体(如表格和图表)的完整段落,而非传统的机器阅读理解。该数据集通过引入多模态文档实体检索框架,旨在提升视觉和语言模型在处理文本主导文档的视觉问答任务中的能力。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务