QUQA
收藏资源简介:
QUQA数据集是由rttl实验室创建的基于《古兰经》的阿拉伯语问答数据集,旨在支持伊斯兰领域的神经检索任务。该数据集包含3382对问答对,经过数据增强后扩展到5385对,涵盖了阿拉伯语和英语的双语数据。数据集的内容主要来源于《古兰经》及其注释,通过Tafseer Ibn Katheer的经文关系生成高质量的领域内数据。该数据集的应用领域为伊斯兰文本的信息检索,旨在提高双语检索模型在伊斯兰文献中的表现,帮助学者和研究人员更高效地检索相关文献。
The QUQA dataset is an Arabic question-answering dataset based on the Quran, created by the RTTL Lab, aiming to support neural retrieval tasks in the Islamic domain. This dataset initially contains 3382 QA pairs, which are expanded to 5385 pairs via data augmentation, covering bilingual data in both Arabic and English. The content of the dataset mainly originates from the Quran and its commentaries, and high-quality in-domain data is generated based on the textual relationships outlined in Tafseer Ibn Katheer. The target application field of this dataset is information retrieval for Islamic texts, with the goal of improving the performance of bilingual retrieval models in Islamic literature and helping scholars and researchers retrieve relevant literature more efficiently.
数据集概述
数据集名称
- QUQA Version 1.0
- HAQA Version 1.0
数据集描述
- QUQA:针对《古兰经》的阿拉伯语问题回答测试集
- HAQA:针对《圣训》的阿拉伯语问题回答测试集
发布日期
- 2023年4月5日
引用信息
@inproceedings{alnefaie2023haqa, title={HAQA and QUQA: Constructing Two Arabic Question-Answering Corpora for the Quran and Hadith}, author={Alnefaie, Sarah and Atwell, Eric and Alsalka, Mohammad Ammar}, booktitle={Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing}, pages={90--97}, year={2023} }
反馈与修正
- 如有关于数据集的修正或评论,请发送邮件至:scsaln@leeds.ac.uk




