yale-nlp/SciDQA
收藏资源简介:
SciDQA是一个专注于科学论文深度阅读理解的数据集,包含2937个问答对。该数据集的问题来源于领域专家的同行评审,答案由论文作者提供,确保了文献的深入审查。数据集通过过滤低质量问题、去上下文化内容、跟踪不同版本的源文档以及包含参考文献来增强质量。问题需要跨图表、方程、附录和补充材料进行推理,并需要多文档推理。该数据集旨在促进复杂科学文本理解的研究。
SciDQA is a dataset for deep reading comprehension over scientific papers, consisting of 2,937 QA pairs. The datasets QA pairs are sourced from peer reviews by domain experts and answers by paper authors, ensuring a thorough examination of the literature. The datasets quality is enhanced through a process that carefully filters out lower quality questions, decontextualizes the content, tracks the source document across different versions, and incorporates a bibliography for multi-document question-answering. Questions in SciDQA necessitate reasoning across figures, tables, equations, appendices, and supplementary materials, and require multi-document reasoning. The dataset is designed to facilitate research on complex scientific text understanding.



