MasterThesisCBS/NorPaca
收藏资源简介:
该数据集是将Alpaca数据集翻译成挪威博克马尔语(Norwegian Bokmål)的版本,并且是基于GPT-4生成的。数据集包含指令、输入、输出和提示等特征,分为训练集和测试集。训练集包含50961个样本,测试集包含1041个样本。数据集的生成提示详细说明了生成指令的要求和格式,包括指令的多样性、语言、类型、长度、输入和输出的要求等。
This dataset is a Norwegian Bokmål translation of the Alpaca dataset, generated using GPT-4. It comprises attributes including instruction, input, output and prompt, and is divided into training and test subsets. The training subset contains 50961 samples, while the test subset has 1041 samples. The dataset's generation prompt details the requirements and formatting specifications for instruction generation, encompassing the diversity, language, type, length of instructions, as well as the requirements for inputs and outputs, among other relevant aspects.
数据集概述
基本信息
- 许可证: cc-by-4.0
- 语言:
- Norwegian Bokmål (nb)
- Norwegian (no)
- 标签: instruction-finetuning
- 任务类别: text-generation
- 数据集名称: NB Alpaca Norwegian Bokmål
数据集结构
-
特征:
- instruction: 字符串类型
- input: 字符串类型
- output: 字符串类型
- prompt: 字符串类型
-
分割:
- 训练集:
- 字节数: 54356020
- 示例数: 50961
- 测试集:
- 字节数: 1113587
- 示例数: 1041
- 训练集:
-
下载大小: 28514339字节
-
数据集大小: 55469607字节



