遇见数据集

strak2005/corpus-ptbr-v1

收藏
Hugging Face2026-05-23 更新2026-05-31 收录
官方服务:

资源简介:

Corpus PT-BR v1 是一个巴西葡萄牙语(pt-br)语料库,包含约840万文档和63亿令牌,专门用于大型语言模型(LLM)的预训练和微调。该数据集结合了经过SBERT质量过滤的真实数据(来自C4和FineWeb2的葡萄牙语子集)和由多种LLM(如Qwen、DeepSeek、Llama等)生成的合成数据层,旨在增强文本的风格、词汇和话语多样性。真实数据部分包括清理和去重后的网络爬取文本,而合成数据部分通过多样化的提示和模型生成,覆盖多种文本风格(如教育文章、访谈、社交媒体帖子等)。数据集还经过去重、长度过滤和标准化处理,以Parquet格式提供,适用于NLP任务如文本生成、填充掩码、文本分类等。

Corpus PT-BR v1 is a Brazilian Portuguese (pt-br) corpus containing approximately 8.4 million documents and 6.3 billion tokens, specially designed for pre-training and fine-tuning of Large Language Models (LLMs). This dataset combines real data filtered by SBERT for quality control (its Portuguese subsets sourced from C4 and FineWeb2) and a layer of synthetic data generated by various LLMs including Qwen, DeepSeek, Llama, and others, aiming to enhance the stylistic, lexical and discourse diversity of the text. The real data component consists of cleaned and deduplicated web-crawled text, while the synthetic data is generated via diverse prompts and multiple models, covering multiple text styles such as educational articles, interviews, social media posts and more. Additionally, the dataset has undergone deduplication, length filtering and standardization processing, and is provided in Parquet format, suitable for NLP tasks including text generation, masked language modeling, text classification and other related tasks.

提供机构:
strak2005
二维码
社区交流群
二维码
科研交流群
商业服务