JDocQA
收藏资源简介:
JDocQA是一个专注于日语文档问题回答的大型数据集,由奈良先端科学技术大学院大学和RIKEN共同创建。该数据集包含5,504个PDF文档,涵盖报告、幻灯片、宣传册和网站等多种格式,共有11,600个问题-答案对。每个问题-答案对都涉及到文档中的文本和视觉元素,如表格或图表,并包含对答案线索的页面引用和边界框标注。JDocQA旨在通过集成文本和视觉信息,评估生成语言模型在实际应用中的问题回答能力,特别强调了不可回答问题的处理,以减少模型产生的幻觉现象。
JDocQA is a large-scale dataset dedicated to Japanese document question answering, co-developed by Nara Institute of Science and Technology and RIKEN. The dataset contains 5,504 PDF documents covering diverse formats including reports, slides, brochures, and web-based materials, with a total of 11,600 question-answer pairs. Each question-answer pair involves both textual and visual elements in the documents, such as tables or charts, and includes page references and bounding box annotations for the answer clues. JDocQA aims to evaluate the question answering capabilities of large language models (LLMs) in real-world applications by integrating textual and visual information, with a particular emphasis on handling unanswerable questions to mitigate hallucinations generated by the models.




