遇见数据集

5-31-2024/cold-cases

收藏
Hugging Face2026-05-18 更新2026-05-31 收录
官方服务:

资源简介:

COLD Cases 是一个包含830万美国法律判决的数据集,涵盖文本和元数据,格式为压缩的parquet文件。该数据集旨在支持开放法律运动,例如Pile of Law和LegalBench项目。关键输入是判例法——法官发表的、具有先例效力的决定,用于解释法律争议和推理过程。美国判例法由CourtListener收集并作为开放数据发布,其通过爬虫从广泛的公共来源聚合数据。COLD Cases重新格式化了CourtListener的批量数据,以便将每个法律判决的语义信息(如多数意见和异议意见的作者和文本、头事项以及实质性元数据)编码为每个判决的单个记录,并移除了无关数据。哈佛图书馆创新实验室作为标准化管理者,维护这个开源管道,以整合预处理判例法的数据工程,使下游机器学习和自然语言处理项目能够使用一致、高质量的案件表示进行法律理解任务。数据集由哈佛图书馆创新实验室与自由法律项目合作准备。

COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. This dataset exists to support the open legal movement exemplified by projects like Pile of Law and LegalBench. A key input to legal understanding projects is caselaw -- the published, precedential decisions of judges deciding legal disputes and explaining their reasoning. United States caselaw is collected and published as open data by CourtListener, which maintains scrapers to aggregate data from a wide range of public sources. COLD Cases reformats CourtListeners bulk data so that all of the semantic information about each legal decision (the authors and text of majority and dissenting opinions; head matter; and substantive metadata) is encoded in a single record per decision, with extraneous data removed. Serving in the traditional role of libraries as a standardization steward, the Harvard Library Innovation Lab is maintaining this open source pipeline to consolidate the data engineering for preprocessing caselaw so downstream machine learning and natural language processing projects can use consistent, high quality representations of cases for legal understanding tasks. Prepared by the Harvard Library Innovation Lab in collaboration with the Free Law Project.

提供机构:
5-31-2024
二维码
社区交流群
二维码
科研交流群
商业服务