Indonesian legislation as data: 90,588 norm units extracted from 378 official regulatory PDFs (1945–2024) with resolvable provenance records
收藏资源简介:
A corpus of 90,588 norm units segmented from 378 official Indonesian regulations (1945–2024), with page- and line-level provenance into the distributed text files. The dataset contains: (1) units.csv — the main table, one row per norm unit, with structural markers (BAB, Paragraf, Pasal, ayat), original and normalised text, IndoBERT token statistics, and a quality flag; (2) documents.csv — the document inventory, with the regulation identity read from each document's title block, the values found in the file name, and a two-axis classification of legal force under Law 12 of 2011 (ladder position under Article 7(1), or delegated force under Article 8); (3) texts.zip — 378 page-delimited plain-text files that the page and line references in units.csv resolve into; (4) scripts.zip — the Python scripts used to build the corpus and compute the quality figures; and (5) checksums.md5 and README.md. Every unit resolves to a document, page, and line range in the distributed text files, so any unit can be verified against its source. The data can be reused for legal NLP tasks such as segmentation, classification, and retrieval, for studies of Indonesian delegated legislation, and as a starting point for converting Indonesian legislation into structured formats such as Akoma Ntoso. Source PDFs are official public documents from Indonesian government legal-documentation portals (JDIH) and are not redistributed here




