LePaRD
收藏资源简介:
LePaRD是一个大规模的法律文本检索数据集,由麻省理工学院和哈佛大学法学院创建。该数据集包含数百万条美国联邦法院的先例引用,旨在促进法律文本预测的研究。数据集的内容包括大量的法律论证上下文和相关的目标文本,主要用于法律领域的文本检索和推理任务。创建过程中,研究者们利用了哈佛的案例法律访问项目(CAP)的数据,通过精确的文本匹配技术提取了引用先例的文本。LePaRD的应用领域主要集中在法律实践,特别是帮助律师和法官减少法律研究的时间和成本,从而扩大司法的可达性。
LePaRD is a large-scale legal text retrieval dataset developed by the Massachusetts Institute of Technology (MIT) and Harvard Law School. It contains millions of U.S. federal court precedent citations, and is designed to facilitate research on legal text prediction. The dataset includes extensive legal argumentation contexts and related target texts, which are mainly used for text retrieval and reasoning tasks in the legal field. During its creation, researchers utilized data from the Harvard Case Law Access Project (CAP), and extracted texts of cited precedents through precise text matching technologies. The application areas of LePaRD mainly focus on legal practice, specifically helping lawyers and judges reduce the time and cost of legal research, thereby expanding access to justice.
LePaRD: A Large-Scale Dataset of Judges Citing Precedents
描述
LePaRD是一个大规模的美国联邦法官引用先例的数据集。该数据集基于数百万专家判决,从中提取了引用先例的引文及其前文上下文。数据集的每一行对应于在特定上下文中使用的先例法律的引用。
数据字段
- passage_id: 每个段落的唯一标识符
- destination_context: 引用前的上下文
- passage_text: 被引用的段落文本
- court: 段落来源的法院
- date: 段落来源的意见书发布的日期
引用
如果使用LePaRD数据集,请引用以下论文: bibtex @article{mahari2023LePaRD, title={LePaRD: A Large-Scale Dataset of Judges Citing Precedents}, author={Mahari, Robert and Stammbach, Dominik and Ash, Elliott and Pentland, AlexSandy}, journal={arXiv preprint}, year={2023} }




