Hybrid-SQuAD
收藏资源简介:
Hybrid-SQuAD是一个用于学术问答(QA)的大型数据集,由汉堡大学和吕讷堡大学等机构创建。该数据集包含10.5K个问题-答案对,利用了DBLP、SemOpenAlex和Wikipedia的文本数据。数据集通过大型语言模型生成,旨在解决学术信息跨异构数据源的问答问题。创建过程中,数据集整合了来自多个数据源的信息,并通过RAG模型进行基线测试,展示了其在学术QA领域的应用潜力。
Hybrid-SQuAD is a large-scale academic question answering (QA) dataset developed by institutions including the University of Hamburg and Leuphana University of Lüneburg. It contains 10.5K question-answer pairs, with textual data sourced from DBLP, SemOpenAlex, and Wikipedia. Generated using large language models (LLMs), this dataset is designed to address the QA challenges of academic information across heterogeneous data sources. During its development, the dataset integrates information from multiple sources, and baseline tests conducted via RAG models have demonstrated its application potential in the academic QA domain.




