synthetic-from-retrieval-tasks-swedish
收藏资源简介:
该数据集的主要目的是用于检索任务的嵌入模型的预训练或后训练。数据集包含100,000个样本,这些样本是通过gemma-2-27b-it模型生成的。每个样本的'prompt'列显示了给LLM的提示,而'response'列显示了LLM的输出。数据集的生成过程遵循了特定论文中描述的方法。
The primary purpose of this dataset is for pre-training or post-training of embedding models for retrieval tasks. This dataset contains 100,000 samples generated using the gemma-2-27b-it model. For each sample, the 'prompt' column displays the input prompt given to the LLM, while the 'response' column shows the output generated by the LLM. The dataset creation process follows the methodology described in a specific academic paper.
数据集概述
数据集名称
ThatsGroes/synthetic-from-retrieval-tasks-swedish
数据集特点
-
特征:
response:字符串类型model:字符串类型prompt:content:字符串类型role:字符串类型
-
数据划分:
- 训练集(train):157,185,409 字节,共 50,000 个样本
-
下载大小:54,152,645 字节
-
数据集大小:157,185,409 字节
-
配置:
default:- 训练集文件路径:data/train-*
-
许可证:MIT
-
任务类别:文本检索(text-retrieval)
-
语言:瑞典语(sv)
数据集描述
本数据集旨在用于预训练或后训练用于检索任务的嵌入模型。数据集包含 100,000 个样本,使用 gemma-2-27b-it 生成。每个样本由从 ThatsGroes/retrieval-tasks-processed 随机抽取的种子任务生成,遵循 此论文 中描述的数据生成过程。计算资源由 Arrow Denmark 和 Nvidia 通过 Danish Data Science Community 赞助。




