arxiv-hard-negatives-cross-encoder
收藏资源简介:
该数据集包含使用交叉编码器生成的硬负样例,用于训练密集检索模型。该数据被用于论文《Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval》中。数据集是Hugging Face收藏的一部分。
This dataset comprises hard negative examples generated by cross-encoders, designed for training dense retrieval models. It has been employed in the paper titled *Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval*, and is part of the Hugging Face collection.
数据集概述
基本信息
- 语言: 英语 (en)
- 数据规模: 1K<n<10K
- 任务类别: 文本排序 (text-ranking)、文本检索 (text-retrieval)
- 许可证: CC-BY-NC-4.0
数据集内容
- 包含通过交叉编码器生成的难负例样本,用于训练密集检索模型。
相关研究
- 数据集用于论文《Dont Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval》。
- 论文链接: https://arxiv.org/abs/2504.21015
引用信息
bibtex @misc{sinha2025dontretrievegenerateprompting, title={Dont Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval}, author={Aarush Sinha}, year={2025}, eprint={2504.21015}, archivePrefix={arXiv}, primaryClass={cs.IR}, url={https://arxiv.org/abs/2504.21015}, }
bibtex @misc{reimers2019sentencebertsentenceembeddingsusing, title={Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks}, author={Nils Reimers and Iryna Gurevych}, year={2019}, eprint={1908.10084}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/1908.10084}, }
其他信息
- 数据集属于Hugging Face集合: arxiv-hard-negatives-68027bbc601ff6cc8eb1f449




