A Dataset Showing a Century of Evolution in the Complexity of the United States Legal Code
收藏资源简介:
We leverage <b>OCR</b> and <b>Generative AI</b> techniques to recover and clean printed historical editions of the Code. This enables computational analysis of federal law even in periods before web-based digital access. The processing pipeline includes:📄 <b>Contents of U.S. Code</b>: Word counts, unique word counts, entropy, scaling exponents, etc.🌲 <b>Hierarchical Structure</b>: Subtitle → Part → Chapter → Section → Subsection...🔗 <b>Cross-Reference Relationships</b>: Title-to-title citation relationshipsFor the small sample of our data, please check out our github repository https://github.com/Dawoon-Jeong0523/uscode-complexity🔍 A sample OCR text page (<code>ocr_processing_gemini</code>) for demonstration🌐 Web-based U.S. Code text from 1994 for structural parsing (<code>Data Set 2</code>)<br>



