遇见数据集

omai-research/bge-multilingual-distillation-dataset

收藏
Hugging Face2026-05-03 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个用于信息检索或排序任务的数据集,包含查询(query)及其对应的正面(pos)和负面(neg)示例列表,每个示例附带有分数(pos_scores和neg_scores)。数据集仅包含训练集,共有1,386,383个示例,总大小约为28.8GB,下载大小约为15.4GB。它适用于训练机器学习模型,以学习查询与相关/不相关内容之间的对比关系。

This dataset is designed for information retrieval or ranking tasks, containing queries (query) along with lists of positive (pos) and negative (neg) examples, each accompanied by scores (pos_scores and neg_scores). The dataset includes only a training split with 1,386,383 examples, totaling approximately 28.8GB in size and a download size of about 15.4GB. It is suitable for training machine learning models to learn contrastive relationships between queries and relevant/irrelevant content.

提供机构:
omai-research
二维码
社区交流群
二维码
科研交流群
商业服务