IL-PCSR
收藏资源简介:
IL-PCSR是一个为印度法律领域量身定制的语料库,旨在解决法律案件中的法律条文检索和先前案例检索问题。该数据集包含6271个案例判决文档、936个法律条文和3183个先前案例,涵盖13个广泛的法律领域。数据集的构建过程涉及从印度Kanoon平台收集20,000份公开可用的英语案例判决书,并通过匿名化处理和事件掩码来防止模型与法律条文和案例标题相关联。IL-PCSR是第一个支持对同一查询并行识别相关法律条文和先前案例的数据集,为法律领域的信息检索模型开发提供了一个共同测试平台。
IL-PCSR is a corpus tailored specifically for the Indian legal domain, developed to address the core challenges of legal provision retrieval and prior case retrieval in legal cases. This dataset contains 6,271 case judgment documents, 936 legal provisions, and 3,183 prior cases, covering 13 broad legal fields. The construction of this dataset involves collecting 20,000 publicly available English case judgments from the Indian Kanoon platform, and applying anonymization and event masking techniques to prevent the model from associating with legal provisions and case titles. IL-PCSR is the first dataset that supports parallel identification of relevant legal provisions and prior cases for the same query, serving as a common testbed for the development of information retrieval models in the legal domain.




