越南语通用文本语料库
收藏资源简介:
本数据集面向越南语大语言模型的训练与迭代,以海量数据驱动模型性能跃升。提供高达12亿条越南语文本,是当前规模最大的越南语训练语源之一。 该体量可支撑训练十亿至百亿参数级别的越南语专用大模型,显著提升其在长文本生成、多轮对话及领域迁移中的稳定性。数据处理过程针对越南语声调符号和复合词边界进行了专门保持,避免预训练中的字符信息丢失。
This dataset is designed for the training and iterative optimization of Vietnamese large language models (LLMs), leveraging massive volumes of data to significantly boost model performance. It contains up to 1.2 billion Vietnamese text instances, making it one of the largest Vietnamese training corpora currently available. This scale enables the training of Vietnamese-specialized LLMs with parameter sizes ranging from 1 billion to 10 billion, greatly improving their stability in long text generation, multi-turn dialogue, and domain adaptation. The data processing pipeline has been specially configured to preserve Vietnamese tone marks and compound word boundaries, preventing the loss of character-level information during pre-training.




