遇见数据集

dataseek/ptbr-gov-legal

收藏
Hugging Face2026-05-12 更新2026-05-31 收录
官方服务:

资源简介:

PT-BR Legal & Government Documents 数据集是 MagTina350m 预训练语料库的一部分,由 Dataseek 发布。它包含 935,000 个巴西法律和政府文件,涵盖联邦/州法律、法院判决、法规和官方通信。该数据集通过合并 eduagarcia/LegalPT_dedup(来自 HuggingFace)和一个 Kaggle 巴西法律程序转储创建,并经过过滤(如 alpha_ratio≥0.65、文本长度≥200 字符、语言检测为葡萄牙语)和去重处理。数据模式包括来源标识、文本内容、字符数、单词数和元数据。数据集大小为 935,685 行,5.96 亿字符,估计 1.32 亿令牌,用于巴西葡萄牙语语言模型的预训练、领域适应和 NLP 研究。许可证主要为 CC0 1.0,但部分数据继承 CC-BY-SA 义务。

PT-BR Legal & Government Documents is part of the MagTina350m pretrain corpus release by Dataseek. It contains 935 K Brazilian legal and government documents, including federal/state laws, court decisions, regulatory acts, and official communications. The dataset is a mixed corpus combining eduagarcia/LegalPT_dedup from HuggingFace with a Kaggle Brazilian-legal-proceedings dump. It has been filtered (alpha_ratio ≥ 0.65, text length ≥ 200 chars, FastText langid=pt) and deduplicated (SHA-1). The schema includes source identifier, text body, character count, word count, and metadata. The dataset has 935,685 rows, 5.96 billion characters, and an estimated 1.32 billion tokens, used for pre-training Brazilian Portuguese language models, domain adaptation, and NLP research. Licensed under CC0 1.0 with some CC-BY-SA obligations.

提供机构:
dataseek
二维码
社区交流群
二维码
科研交流群
商业服务