envyr/natural-questions-hard-negatives_gemma_MarginMSE_instructions_STATISTICAL
收藏数据链接:
官方服务:
资源简介:
该数据集是一个用于训练任务的数据集,包含三元组文本数据,每个样本由anchor(锚点文本)、positive(正例文本)和negative(负例文本)组成,并附带一个浮点型标签(label)。数据集仅包含训练分割,共有484,310个样本,总大小约为678.7 MB,适用于对比学习或相似性学习等NLP任务。
This dataset is designed for training tasks and consists of triplet text data, where each sample includes an anchor text, a positive text, and a negative text, along with a float label. It contains only a training split with 484,310 examples and a total size of approximately 678.7 MB, suitable for NLP tasks such as contrastive learning or similarity learning.
提供机构:
envyr


