遇见数据集

TeraflopAI/Caselaw_Access_Project

收藏
Hugging Face2024-03-16 更新2024-04-19 收录
官方服务:

资源简介:

--- license: cc0-1.0 task_categories: - text-generation language: - en tags: - legal - law - caselaw pretty_name: Caselaw Access Project size_categories: - 1M<n<10M --- <img src="https://huggingface.co/datasets/TeraflopAI/Caselaw_Access_project/resolve/main/cap.png" width="800"> # The Caselaw Access Project In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/ Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here: https://case.law/docs/ Learn more about the Caselaw Access Project and all of the phenomenal work done by Jack Cushman, Greg Leppert, and Matteo Cargnelutti here: https://case.law/about/ Watch a live stream of the data release here: https://lil.law.harvard.edu/about/cap-celebration/stream # Post-processing Teraflop AI is excited to help support the Caselaw Access Project and Harvard Library Innovation Lab, in the release of over 6.6 million state and federal court decisions published throughout U.S. history. It is important to democratize fair access to data to the public, legal community, and researchers. This is a processed and cleaned version of the original CAP data. During the digitization of these texts, there were erroneous OCR errors that occurred. We worked to post-process each of the texts for model training to fix encoding, normalization, repetition, redundancy, parsing, and formatting. Teraflop AI’s data engine allows for the massively parallel processing of web-scale datasets into cleaned text form. Our one-click deployment allowed us to easily split the computation between 1000s of nodes on our managed infrastructure. # Licensing Information The Caselaw Access Project dataset is licensed under the [CC0 License](https://creativecommons.org/public-domain/cc0/). # Citation Information ``` The President and Fellows of Harvard University. "Caselaw Access Project." 2024, https://case.law/ ``` ``` @misc{ccap, title={Cleaned Caselaw Access Project}, author={Enrico Shippole, Aran Komatsuzaki}, howpublished{\url{https://huggingface.co/datasets/TeraflopAI/Caselaw_Access_Project}}, year={2024} } ```

The Caselaw Access Project dataset was digitized through a collaboration between the Harvard Law School Library and Ravel Law. It encompasses over 40 million U.S. court decisions, covering 6.7 million cases spanning a 360-year period. Teraflop AI conducted post-processing on the dataset, including fixing OCR errors, as well as encoding, normalization, deduplication, redundancy removal, parsing, and formatting. The dataset is intended to provide equitable data access for the general public, legal community, and researchers. It is released under the CC0 license, allowing for widespread and unrestricted usage.

提供机构:
TeraflopAI
原始信息汇总

数据集概述

数据集来源

  • 合作方:Ravel Law 与 Harvard Law Library

数据集内容

  • 包含超过4000万份美国法院判决
  • 涵盖670万案件
  • 时间跨度:过去360年

数据集访问方式

搜集汇总
数据集介绍
TeraflopAI/Caselaw_Access_Project 数据集图片
背景与挑战
背景概述
该数据集是Teraflop AI处理后的Caselaw Access Project版本,包含美国法院的超过40百万个决策和6.7百万个案件,时间跨度360年,适用于法律研究和文本生成任务。数据集经过清理以修复OCR错误和格式化问题,采用parquet格式和CC0许可证,便于公众和研究人员访问。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务