PerCQA
收藏资源简介:
PerCQA是第一个用于社区问答的波斯数据集。 社区问答 (CQA) 论坛为许多现实生活中的问题提供答案。由于规模庞大,这些论坛在机器学习研究人员中非常受欢迎。自动选择答案,答案排名,问题检索,专家发现和事实检查是使用CQA数据执行的示例学习任务。 在本文中,我们介绍了PerCQA,这是CQA的第一个波斯数据集。此数据集包含从最著名的波斯论坛抓取的问题和答案。数据采集后,我们在迭代过程中提供严格的注释指南,然后以SemEvalCQA格式对问答对进行注释。 PerCQA包含989个问题和21,915带注释的答案。我们公开提供PerCQA,以鼓励对波斯CQA进行更多研究。我们还通过使用单语言和多语言的预训练语言模型,为PerCQA中的答案选择任务建立了强大的基准。
PerCQA is the first Persian dataset for community question answering (CQA). Community question answering (CQA) forums provide answers to many real-world problems. Due to their large scale, these forums are extremely popular among machine learning researchers. Typical learning tasks performed using CQA data include automatic answer selection, answer ranking, question retrieval, expert discovery, and fact checking. In this paper, we introduce PerCQA, the first Persian dataset for CQA. This dataset consists of questions and answers scraped from the most well-known Persian forums. After data collection, we developed strict annotation guidelines through an iterative process, and then annotated the question-answer pairs in the SemEval CQA format. PerCQA contains 989 questions and 21,915 annotated answers. We publicly release PerCQA to encourage further research on Persian CQA. We also establish strong benchmarks for the answer selection task on PerCQA using both monolingual and multilingual pretrained language models.




