redis/langcache-triplets-v3
收藏资源简介:
Redis LangCache Triplets Dataset v3是一个用于对比学习的大规模三元组数据集,包含锚点句子、语义相似的正面句子和不相似的负面句子。该数据集来源于Redis LangCache Sentence Pairs v3,结合了多个高质量的转述语料库。数据集主要用于训练句子编码器,适用于语义检索和重新排序等任务。数据集包含约8200万个三元组,全部为英文,采用Apache-2.0许可证。
Redis LangCache Triplets Dataset v3 is a large-scale triplet dataset for training sentence encoders using contrastive learning. Each example contains an anchor sentence, a semantically similar positive sentence, and a dissimilar negative sentence. The triplets are generated from the LangCache Sentence Pairs v3 dataset, which combines multiple high-quality paraphrase corpora. The dataset is primarily used for training sentence encoders and is suitable for tasks like semantic retrieval and re-ranking. It contains approximately 82 million triplets, all in English, and is licensed under Apache-2.0.



