遇见数据集

Madras1/corpus-ptbr-v2

收藏
Hugging Face2026-05-24 更新2026-05-31 收录
官方服务:

资源简介:

Corpus PT-BR v2 是一个面向大型语言模型(LLMs)预训练、继续预训练和微调的巴西葡萄牙语语料库。该版本基于 `Madras1/corpus-ptbr-v1` 的框架和流程,并主要通过 Mistral 模型扩展了合成数据层。数据集包含真实数据(来自公共来源如 Common Crawl 和 FineWeb2 的清洗过滤数据)和合成数据(由 Mistral 等模型生成),总计约 877 万文档、54.3 亿单词和估计 70.6 亿令牌。数据以 Parquet 格式提供,支持流式加载,并遵循 ODC-By 1.0 许可证。

Corpus PT-BR v2 is a Brazilian Portuguese corpus aimed at pre-training, continued pre-training, and fine-tuning of LLMs. This version maintains the base and general pipeline of `Madras1/corpus-ptbr-v1`, with an additional expansion of the synthetic layer generated primarily by Mistral models. The dataset includes real data (cleaned and filtered from public sources like Common Crawl and FineWeb2) and synthetic data (generated by models such as Mistral), totaling approximately 8.77 million documents, 5.43 billion words, and an estimated 7.06 billion tokens. It is provided in Parquet format, supports streaming loading, and is licensed under ODC-By 1.0.

提供机构:
Madras1
二维码
社区交流群
二维码
科研交流群
商业服务