Itau-Unibanco/Aurora-PT
收藏资源简介:
Aurora-PT是一个为葡萄牙语大型语言模型训练设计的精选语料库。现有研究大多集中在英语和汉语等高资源语言,并致力于开发多语言语料库,但开发低资源语言的大型数据集需求迫切。该工作旨在缩小这一差距,促进葡萄牙语最先进研究的发展。据我们所知,Aurora-PT是最大的公开单语葡萄牙语语料库,超越了所有先前资源。最终语料库是Itaú Unibanco和Itaú-Unibanco科学技术研究所(ICTi)研究人员共同工作的成果,其构建细节在论文“NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus”(PROPOR 2026)中描述。数据集开发者:Itaú Unibanco和Itaú-Unibanco科学技术研究所(ICTi)。数据集构建架构:受FineWeb v1的流程启发,但关键区别在于聚合了九个葡萄牙语或多语言数据集(CC100、mOSCAR、Aya、FineWeb v2、Blogset-br、Aroeira、mC4、Wikipedia和HPLT 2.0),而非仅依赖Common Crawl转储。预处理流程包括多个阶段:使用GlotLID进行语言过滤(葡萄牙语阈值0.799)、通过MinHash进行去重(14个桶中的112个哈希)、基于C4的质量启发式规则和FineWeb特定质量规则。所有处理步骤均使用Hugging Face的DataTrove库实现。支持语言:巴西葡萄牙语(PT-BR)和葡萄牙语(PT-PT)。数据集发布日期:2026年。状态:Aurora-PT包含约855 GB(未压缩)、2.649亿个文档和3310亿个令牌(GPT-2分词)/2260亿个令牌(语料训练分词器)。
Aurora-PT is a curated corpus designed for training large language models in the Portuguese language. Most existing research focuses on high-resource languages like English and Chinese, with considerable efforts made to develop multilingual corpora. However, there is a pressing need to develop large datasets for lower-resource languages. This work aims to make this gap smaller, contributing to the development of state-of-the-art research in Portuguese. Aurora-PT is, to our knowledge, the largest openly available monolingual Portuguese corpus, surpassing all previous resources. The final corpus is a result of the combined work of researchers at Itaú Unibanco and the Instituto de Ciência e Tecnologia Itaú-Unibanco (ICTi), and the details about how it was made are described in the paper "NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus" (PROPOR 2026). Dataset developer: Itaú Unibanco and Instituto de Ciência e Tecnologia Itaú-Unibanco (ICTi). Dataset Creation Architecture: The construction of Aurora-PT was inspired by the pipeline used in FineWeb v1, with a key distinction: instead of relying solely on Common Crawl dumps, nine Portuguese or multilingual datasets annotated by language were aggregated (CC100, mOSCAR, Aya, FineWeb v2, Blogset-br, Aroeira, mC4, Wikipedia, and HPLT 2.0). The preprocessing pipeline comprised several stages: language filtering using GlotLID (threshold 0.799 for Portuguese), deduplication via MinHash (112 hashes across 14 buckets), C4-based quality heuristics, and FineWeb-specific quality rules. All processing steps were implemented using the DataTrove library from Hugging Face. Supported languages: Brazilian Portuguese (PT-BR) and Portuguese (PT-PT). Dataset Release Date: 2026. Status: Aurora-PT contains approximately 855 GB (uncompressed), 264.9 million documents, and 331 billion tokens (GPT-2 tokenization) / 226 billion tokens (corpus-trained tokenizer).




