InPars+
收藏资源简介:
InPars+数据集是由阿姆斯特丹大学的研究团队开发的,用于信息检索系统的人工合成数据生成。该数据集通过InPars Toolkit生成,这是一个可重复的、端到端的人工合成数据生成框架,利用大型语言模型(LLM)进行训练数据生成。数据集的大小、数据量和Tokens数等信息在论文中没有明确提及。InPars+数据集旨在解决信息检索模型训练数据不足的问题,通过合成相关查询来提高模型的检索性能。
The InPars+ dataset was developed by the research team from the University of Amsterdam for synthetic data generation in information retrieval systems. This dataset is generated via the InPars Toolkit, a reproducible, end-to-end synthetic data generation framework that leverages Large Language Models (LLMs) to produce training data. Specific details including the dataset size, total data volume, and number of tokens are not explicitly mentioned in the associated paper. The InPars+ dataset aims to address the shortage of training data for information retrieval models, and improves the retrieval performance of such models by generating relevant synthetic queries.




