alina0195/ro-msmarco
收藏资源简介:
该数据集是一个用于对比学习或三元组任务的大规模文本数据集,包含三个主要特征:anchor(锚点)、positive(正例)和negative(负例),均为字符串类型。数据集仅提供训练分割,包含约1166万个示例,总大小约为9.7 GB,下载大小约为2.5 GB。数据以文件形式存储,路径为data/train-*。数据集可能适用于自然语言处理任务,如文本相似度计算或嵌入学习,但README未明确说明具体应用领域或数据来源。
This dataset is a large-scale text dataset designed for contrastive learning or triplet tasks, featuring three main components: anchor, positive, and negative, all of string type. It includes only a training split with approximately 11.66 million examples, a total size of about 9.7 GB, and a download size of about 2.5 GB. The data is stored in files with the path data/train-*. The dataset is likely suitable for natural language processing tasks such as text similarity calculation or embedding learning, but the README does not specify the exact application domain or data source.



