GeAR
收藏资源简介:
GeAR数据集由微软公司创建,旨在支持生成增强检索(GeAR)模型的训练。该数据集包含580万条数据,主要用于问题回答检索(QAR)和相关信息检索(RIR)任务。数据来源于高质量维基百科文档,通过大语言模型(LLM)生成查询和文档的细粒度信息单元,并经过去重和相关性过滤处理。数据集的应用领域包括文档检索、细粒度信息定位和信息生成,旨在提升检索系统对复杂文本的细粒度语义理解能力。
The GeAR dataset was created by Microsoft to support the training of Generative-Augmented Retrieval (GeAR) models. It contains 5.8 million data entries, and is primarily used for Question Answering Retrieval (QAR) and Relevant Information Retrieval (RIR) tasks. The dataset is sourced from high-quality Wikipedia documents, where fine-grained information units of queries and documents are generated by Large Language Models (LLMs), and then subjected to deduplication and relevance filtering processing. Its application areas include document retrieval, fine-grained information localization, and information generation, with the goal of enhancing the fine-grained semantic comprehension capabilities of retrieval systems for complex texts.




