nb-gpt-gemma3-12b-base-aurora-2604-open-pretrain
收藏资源简介:
aurora-open-2604 是一个用于挪威语预训练的开源数据集,专注于挪威语及其两种官方书面变体:博克马尔语(Bokmål)和尼诺斯克语(Nynorsk)。该数据集旨在支持大规模语言模型的预训练任务,特别为适应挪威语的语言特点和文化背景而构建。基于该数据集训练的模型(如nb-gpt-gemma3-12b-base-aurora-2604-open-pretrain)使用了3000亿个token进行预训练,表明数据集具有相当大的规模。它适用于挪威语的自然语言处理基础模型开发、语言理解与生成任务,以及多方言挪威语的语言模型研究。
aurora-open-2604 is an open-source dataset for Norwegian language pretraining, focusing on Norwegian and its two official written variants: Bokmål and Nynorsk. The dataset is designed to support large-scale language model pretraining tasks, specifically built to adapt to the linguistic characteristics and cultural context of Norwegian. Models trained on this dataset, such as nb-gpt-gemma3-12b-base-aurora-2604-open-pretrain, use 300 billion tokens for pretraining, indicating that the dataset is of substantial size. It is suitable for Norwegian natural language processing foundation model development, language understanding and generation tasks, as well as research on multi-dialect Norwegian language models.
- 数据集名称: nb-gpt-gemma3-12b-base-aurora-2604-open-pretrain
- 模型基础: google/gemma-3-12b-pt
- 训练数据集: NbAiLab/aurora-open-2604
- 预训练规模: 300B tokens
- 许可协议: 其他 (license: other)
- 语言: 挪威语,包含书面挪威语 (nb) 和新挪威语 (nn)
- 数据集标签: borealis, gemma3, norwegian, norwegian-bokmal, norwegian-nynorsk, pretraining, pretrain, base
- 数据集用途: 基础模型预训练




