遇见数据集

KhaledReda/pairs_with_scores_v45

收藏
Hugging Face2026-05-14 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含句子对及其相似度分数,用于文本相似度或语义相似度任务。每个样本包含两个句子(sentence1和sentence2)以及一个浮点数分数(score),表示它们之间的相似度。训练集包含约2110万条样本,评估集包含约10.6万条样本。

This dataset contains sentence pairs with similarity scores, used for text similarity or semantic similarity tasks. Each sample consists of two sentences (sentence1 and sentence2) and a floating-point score indicating their similarity. The training set contains approximately 21.1 million samples, and the evaluation set contains approximately 106,000 samples.

提供机构:
KhaledReda
二维码
社区交流群
二维码
科研交流群
商业服务