IIRC
收藏资源简介:
IIRC是一个由华盛顿大学和艾伦人工智能研究所合作创建的数据集,包含超过13000个基于英文维基百科段落的不完整信息阅读理解问题。该数据集的特点是每个问题都需要从原始段落及其链接的文档中获取缺失信息才能解答,这要求模型具备跨文档的复杂推理能力。数据集的创建过程通过众包完成,确保了问题的自然性和信息寻求性。IIRC主要用于评估和提升机器阅读理解系统在处理不完整信息时的性能,特别是在识别和检索缺失信息以及综合这些信息以回答问题方面的能力。
IIRC is a dataset co-created by the University of Washington and the Allen Institute for Artificial Intelligence, containing over 13,000 incomplete-information reading comprehension questions based on English Wikipedia paragraphs. A core characteristic of this dataset is that every question requires obtaining missing information from both the original paragraph and its linked documents to be solved, which demands the model to possess cross-document complex reasoning capabilities. The dataset was developed via crowdsourcing, ensuring the naturalness and information-seeking nature of the questions. IIRC is primarily used to evaluate and improve the performance of machine reading comprehension systems when handling incomplete information, especially their abilities to identify and retrieve missing information and synthesize such information to answer questions.




