proxectonos/corpus_dominio_legal_administrativo_tmp
收藏资源简介:
该数据集是一个法律-行政领域语料库,汇集了来自加利西亚官方公报和机构期刊的正式文本,代表了法律-行政语言的形式化、规范化和行政使用。文本包括完整、结构化且带有相关元数据的文档。这是语料库的临时更新版本,其中大幅扩展了来自加利西亚官方期刊的加利西亚语子集,并对省级公报应用了新的匿名化和伪名化处理。数据集包含三个独立子语料库,分别对应加利西亚的三个参考机构公报:拉科鲁尼亚省官方公报(拉科鲁尼亚省议会)、蓬特韦德拉省官方公报(蓬特韦德拉省议会)和加利西亚官方期刊(加利西亚政府)。原始文档以HTML、XHTML或XML格式提供,这些是提取的主要来源。数据以JSONL格式存储,每个文档对应一行JSON对象,字段包括文档标识符、完整行政文本、来源和公报、发布日期、检测到的语言以及可用的行政元数据。数据集适用于特定领域的语言建模、法律和行政语言的持续预训练、加利西亚语和西班牙语行政和法律语料分析、官方行政文本的双语或跨语言实验、机构文本的信息提取和文档分析实验,以及公共行政语言模型的开发和评估。数据集基于Creative Commons Attribution 4.0 International (CC BY 4.0)许可证分发,并得到欧盟NextGenerationEU的资助。
The dataset is a legal-administrative domain corpus that gathers official texts from bulletins and institutional journals of Galicia, representative of the formal, normative, and administrative use of legal-administrative language. The texts included correspond to complete, structured documents with associated metadata. This is a temporary updated version of the corpus, which incorporates a substantial expansion of the Galician subset from the Diario Oficial de Galicia and a new phase of anonymization and pseudonymization applied to provincial bulletins. The dataset consists of three independent subcorpora, each associated with a reference institutional bulletin in Galicia: the Boletín Oficial de la Provincia de A Coruña (Diputación de A Coruña), the Boletín Oficial de la Provincia de Pontevedra (Diputación de Pontevedra), and the Diario Oficial de Galicia (Xunta de Galicia). The original documents were available in HTML, XHTML, or XML formats, which constitute the primary source of extraction. Data is stored in JSONL format, with one JSON object per line per document, including fields such as document identifier, full administrative text, source and bulletin, publication date, detected language, and administrative metadata when available. The dataset is intended for domain-specific language modeling, continuous pretraining on legal and administrative language, corpus analysis of administrative and legal language in Galician and Spanish, bilingual or cross-lingual experiments with official administrative text, information extraction and document analysis experiments on institutional texts, and development and evaluation of models adapted to public administration language. It is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license and funded by the EU NextGenerationEU.




