遇见数据集

envyr/natural-questions-hard-negatives_gemma_MarginMSE_instructions_STATISTICAL

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个用于训练任务的数据集,包含三元组文本数据,每个样本由anchor(锚点文本)、positive(正例文本)和negative(负例文本)组成,并附带一个浮点型标签(label)。数据集仅包含训练分割,共有484,310个样本,总大小约为678.7 MB,适用于对比学习或相似性学习等NLP任务。

This dataset is designed for training tasks and consists of triplet text data, where each sample includes an anchor text, a positive text, and a negative text, along with a float label. It contains only a training split with 484,310 examples and a total size of approximately 678.7 MB, suitable for NLP tasks such as contrastive learning or similarity learning.

提供机构:
envyr
二维码
社区交流群
二维码
科研交流群
商业服务