pretrain-bm-zh
收藏资源简介:
数据集 pretrain-bm-zh 是一个用于机器翻译任务的平行语料库。它基于 MiniMind 项目的 pretrain_split_zh 中文预训练文本,通过 Google TranslateGemma 27B 模型批量翻译成马来语(Bahasa Melayu)。目前包含超过 239 万条翻译对(行),总大小约为 3GB,翻译工作仍在进行中,目标规模约为 777 万条。该数据集旨在为马来语-中文的双语自然语言处理任务,如翻译模型训练或跨语言理解,提供大规模、自动生成的训练或评估资源。数据以文本行形式组织,每行对应一个源语言(中文)句子及其目标语言(马来语)翻译。
The dataset pretrain-bm-zh is a parallel corpus designed for machine translation tasks. Derived from the Chinese pretraining text pretrain_split_zh of the MiniMind project, it was batch-translated into Bahasa Melayu via the Google TranslateGemma 27B model. Currently, it comprises over 2.39 million translation pairs (lines) with a total size of approximately 3 GB. The translation process is still underway, with a target scale of around 7.77 million pairs. This dataset is intended to offer large-scale, automatically generated training and evaluation resources for Malay-Chinese bilingual natural language processing tasks, including translation model training and cross-lingual understanding. The data is structured in text lines, where each line contains a source-language (Chinese) sentence paired with its target-language (Bahasa Melayu) translation.
数据集概述
- 名称: pretrain-bm-zh
- 所有者: khursanirevo
- 语言: 马来语 (ms)、中文 (zh)
- 任务类别: 翻译 (translation)
- 标签: malay、bahasa-melayu、pretrain、minimind
- 数据集大小: 1M 至 10M 行
数据集描述
该数据集是将 MiniMind 项目的中文预训练文本 pretrain_split_zh 翻译为马来语的版本。翻译工作由 Google TranslateGemma 27B 模型(使用 vLLM 实现张量并行度 TP=2,BF16 精度)执行。
当前状态
- 已翻译行数:2,619,679 / 7,775,505(完成率 33.69%)
- 文件大小:3299.6 MB
- 最后推送时间:2026-06-18 04:10:35 +08
- 翻译持续进行中,此文件在翻译运行期间每小时更新一次。




