Cyyyx/LCQMC
收藏资源简介:
LCQMC是一个中文文本相关性数据集,专门用于语义相似性评分任务。该数据集包含句子对(sentence1和sentence2)以及一个整数分数(score),用于表示两个句子之间的相关程度。数据集分为训练集(238,766个示例)、验证集(8,802个示例)和测试集(12,500个示例),总大小约为21MB。它支持自然语言处理中的句子相似性评估,是Massive Text Embedding Benchmark(MTEB)的一部分,适用于训练和测试中文嵌入模型。
LCQMC is a Chinese dataset for textual relatedness, designed for semantic similarity scoring tasks. It consists of sentence pairs (sentence1 and sentence2) along with an integer score indicating the degree of relatedness between the sentences. The dataset is split into train (238,766 examples), validation (8,802 examples), and test (12,500 examples) sets, with a total size of approximately 21MB. It is used for evaluating sentence similarity in natural language processing and is part of the Massive Text Embedding Benchmark (MTEB), suitable for training and testing Chinese embedding models.



