Y-NQ
收藏资源简介:
Y-NQ是一个用于开放书籍阅读理解和文本生成的英-约鲁巴语评估数据集,旨在评估模型在高资源语言(英语)和低资源语言(约鲁巴语)中的表现。数据集包含358个问题和答案,涉及338篇英语文档和208篇约鲁巴语文档,平均文档长度分别为10,000字和430字。数据集的创建过程包括从NQ数据集中筛选问题,并通过人工注释确保答案的准确性。该数据集主要用于评估大型语言模型在不同语言环境下的阅读理解能力,特别是探索英语模型的能力是否能扩展到约鲁巴语。
Y-NQ is an English-Yoruba evaluation dataset for open-book reading comprehension and text generation, designed to assess model performance in both high-resource language (English) and low-resource language (Yoruba). The dataset contains 358 question-answer pairs, involving 338 English documents and 208 Yoruba documents, with average document lengths of 10,000 words and 430 words respectively. The dataset construction process includes screening questions from the NQ dataset and ensuring answer accuracy via manual annotation. This dataset is primarily used to evaluate the reading comprehension capabilities of large language models across different linguistic contexts, particularly to explore whether the capabilities of English models can be extended to Yoruba.

- 1Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text GenerationMeta的FAIR · 2024年



