NorOpenBookQA, NorCommonSenseQA, NorTruthfulQA, NRK-Quiz-QA
收藏资源简介:
该论文介绍了一套新的挪威语问答数据集,包括NorOpenBookQA、NorCommonSenseQA、NorTruthfulQA和NRK-Quiz-QA。这些数据集由奥斯陆大学的研究团队创建,涵盖了挪威语的两种书面标准——Bokmål和Nynorsk。数据集包含超过10,500个问题-答案对,涉及世界知识、常识推理、真实性和挪威相关知识。数据集的创建过程包括手动翻译和本地化英语数据集,并生成新的挪威语示例。这些数据集旨在评估语言模型在挪威语理解和生成方面的能力,特别是在多领域知识和常识推理方面的表现。数据集的应用领域包括自然语言处理、问答系统和语言模型评估,旨在解决挪威语资源匮乏的问题。
This paper introduces a new suite of Norwegian question-answering datasets, namely NorOpenBookQA, NorCommonSenseQA, NorTruthfulQA, and NRK-Quiz-QA. These datasets were developed by the research team at the University of Oslo, covering two standard written varieties of Norwegian: Bokmål and Nynorsk. The suite contains over 10,500 question-answer pairs, spanning world knowledge, commonsense reasoning, truthfulness, and Norway-related knowledge. The creation process of these datasets involves manual translation and localization of existing English datasets, as well as the generation of novel Norwegian examples. These datasets are intended to evaluate the capabilities of language models in Norwegian language understanding and generation, especially their performance across multi-domain knowledge and commonsense reasoning tasks. Their application areas include natural language processing, question answering systems, and language model evaluation, with the goal of addressing the scarcity of Norwegian language resources.




