PUGG
收藏资源简介:
PUGG数据集由弗罗茨瓦夫理工大学开发,是首个波兰语知识库问答(KBQA)资源,涵盖KBQA、机器阅读理解(MRC)和信息检索(IR)三个任务。该数据集包含自然发生的波兰语事实问题,并通过半自动化流程创建,减少了人工标注的工作量。数据集的创建过程中未使用翻译,确保了数据的自然性。PUGG数据集主要用于解决低资源语言环境下AI和自然语言处理中的问答系统问题。
The PUGG dataset, developed by Wrocław University of Science and Technology, is the first Polish-language knowledge base question answering (KBQA) resource covering three tasks: KBQA, machine reading comprehension (MRC), and information retrieval (IR). This dataset contains naturally occurring Polish factual questions and was constructed via a semi-automated pipeline, which reduces the workload of manual annotation. No translation was used during the dataset creation process, ensuring the naturalness of the data. The PUGG dataset is primarily designed to address question answering system challenges in AI and natural language processing in low-resource language settings.
PUGG: KBQA, MRC, IR Dataset for Polish
数据集概述
基本信息
- 语言: 波兰语 (pl)
- 许可证: CC BY-SA 4.0
- 多语言性: 单语种 (monolingual)
- 数据量: 1K<n<10K 和 10K<n<100K
- 数据来源: 原始数据 (original)
- 任务类别:
- 问答 (question-answering)
- 文本检索 (text-retrieval)
- 任务ID:
- 抽取式问答 (extractive-qa)
- 文档检索 (document-retrieval)
数据集配置
- kbqa_all:
- 训练集:
kbqa/*/train.jsonl - 测试集:
kbqa/*/test.jsonl
- 训练集:
- kbqa_natural:
- 训练集:
kbqa/natural/train.jsonl - 测试集:
kbqa/natural/test.jsonl
- 训练集:
- kbqa_template-based:
- 训练集:
kbqa/template-based/train.jsonl - 测试集:
kbqa/template-based/test.jsonl
- 训练集:
- mrc:
- 训练集:
mrc/train.jsonl - 测试集:
mrc/test.jsonl
- 训练集:
- ir_corpus:
- 测试集:
ir/corpus.jsonl
- 测试集:
- ir_queries:
- 测试集:
ir/queries.jsonl
- 测试集:
- ir_qrels:
- 测试集:
ir/qrels/test.jsonl
- 测试集:
标签
- 知识图谱 (knowledge graph)
- KBQA
- 维基百科 (wikipedia)
- 维基数据 (wikidata)




