us-caselaw-pa
收藏资源简介:
Pennsylvania Case Law 是一个包含347,902份宾夕法尼亚州上诉法院意见全文的数据集,来源于 Free Law Project 的 CourtListener 批量导出(截至2026年6月30日)。覆盖三个法院(pa、pacommwct、pasuperct),包括主意见和单独意见,排除字符数小于1的文档。数据格式采用 docketx record v1,每行一条意见记录,字段包括 id、doc_type、jurisdiction、title、text、source、license、retrieved_at、citation、court、date 和 extra(包含 CourtListener 意见/案件 ID 和意见类型)。数据集中,208,631(60.0%)个文档超过2000字符,63,389(18.2%)个在200-1999字符之间,75,882(21.8%)个小于200字符(如拒绝调卷令、单行命令和判决录入等真实记录,保留以保持完整公共记录)。该数据集基于公共领域的司法意见(无版权),以 CC0 1.0 许可发布,无任何人工标注。适用于文本检索、文本生成、法律 NLP、检索增强生成(RAG)以及法律研究等任务。
Pennsylvania Case Law is a dataset containing 347,902 full-text opinions from the Pennsylvania appellate courts, sourced from the Free Law Projects CourtListener bulk export (as of June 30, 2026). It covers three courts (pa, pacommwct, pasuperct), includes majority and separate opinions, and excludes documents with fewer than 1 character. The data format follows docketx record v1, with each row representing one opinion, including fields: id, doc_type, jurisdiction, title, text, source, license, retrieved_at, citation, court, date, and extra (containing CourtListener opinion/case IDs and opinion type). In the dataset, 208,631 (60.0%) documents exceed 2000 characters, 63,389 (18.2%) have 200-1999 characters, and 75,882 (21.8%) have fewer than 200 characters (such as denials of certiorari, per curiam orders, and judgment entries, which are real records retained to maintain a complete public record). The dataset is based on public domain judicial opinions (no copyright) and released under the CC0 1.0 license, with no human annotations. It is suitable for tasks such as text retrieval, text generation, legal NLP, retrieval-augmented generation (RAG), and legal research.





