遇见数据集

KhaledReda/pairs_with_scores_v44

收藏
Hugging Face2026-05-08 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个用于句子对任务的文本数据集,包含两个句子字段(sentence1和sentence2)和一个分数字段(score),分数可能表示句子之间的相似度或相关性。数据集分为训练集(train)和评估集(eval),训练集有29,220,823个示例,评估集有146,839个示例,总大小约为6.7GB。数据以字符串和浮点类型存储,适用于自然语言处理任务,如文本匹配、语义相似度计算或自然语言推理。

This dataset is a text dataset for sentence pair tasks, containing two sentence fields (sentence1 and sentence2) and a score field, which likely indicates similarity or relevance between the sentences. The dataset is split into a training set (train) and an evaluation set (eval), with 29,220,823 examples in the training set and 146,839 examples in the evaluation set, totaling approximately 6.7GB in size. The data is stored as strings and floats, making it suitable for natural language processing tasks such as text matching, semantic similarity calculation, or natural language inference.

提供机构:
KhaledReda
二维码
社区交流群
二维码
科研交流群
商业服务