SCIDQA
收藏资源简介:
SCIDQA数据集是由耶鲁大学和Allen Institute for AI共同创建的一个深度阅读理解数据集,专门用于评估语言模型对科学论文的理解能力。该数据集包含2937个问答对,来源于OpenReview平台上的同行评审,确保了问题和答案的高质量。数据集的创建过程包括从PDF转换、正则表达式过滤到LLM提取问答对,并通过领域专家的手动标注和编辑来确保数据质量。SCIDQA数据集的应用领域主要集中在科学文本理解,旨在解决复杂科学文本的深度理解和推理问题。
The SCIDQA dataset, co-created by Yale University and the Allen Institute for AI, is a deep reading comprehension dataset specifically designed to evaluate language models' ability to comprehend scientific papers. It comprises 2,937 question-answer pairs sourced from peer reviews on the OpenReview platform, which guarantees the high quality of both the questions and their corresponding answers. The dataset construction workflow includes PDF conversion, regular expression filtering, question-answer pair extraction using large language models (LLMs), as well as manual annotation and editing by domain experts to ensure data quality. The SCIDQA dataset is primarily applied in the field of scientific text understanding, aiming to address challenges in deep comprehension and reasoning over complex scientific texts.




