KhaledReda/pairs_with_scores_v43
收藏资源简介:
该数据集是一个用于评估句子对相似度的数据集,包含两个文本字段(sentence1和sentence2)和一个浮点分数(score),分数可能表示句子之间的语义相似度或相关性。数据集分为训练集和评估集,训练集约有4016万个示例,评估集约有20万个示例,总大小约73.4亿字节。它适用于自然语言处理任务,如文本匹配、相似度计算或自然语言推理。
This dataset is designed for evaluating sentence pair similarity. It contains two text fields (sentence1 and sentence2) and a floating-point score (score), which indicates the semantic similarity or relevance between the two sentences. The dataset is split into a training set and an evaluation set, with approximately 40.16 million examples in the training set and 200,000 examples in the evaluation set, and the total size is about 7.34 billion bytes. It is suitable for natural language processing tasks such as text matching, similarity calculation and natural language inference.



