KanoonGPT/indian-legal-documents
收藏资源简介:
该数据集名为印度法律文件,是KanoonGPT开放法律数据倡议的一部分,旨在为AI、搜索和法律研究提供结构化的印度法定和法律文件数据。数据集包含35,851个印度法律文件,涵盖法案、规则、通知、通告、命令、法规、指南、条例等多种文档类型,覆盖中央和多个邦(如马哈拉施特拉邦、北方邦等)的司法管辖区。时间覆盖范围从1970年到2026年,包含解析日期和未解析日期的文件。数据集以AI就绪、表格化格式存储,支持文本生成、问答、特征提取、文本检索、文本排名和摘要等多种NLP任务。数据模式包括稳定文档标识符、文档标题、文档类型、司法管辖区、颁发机构、颁发日期、简短描述和清理后的文本字段。数据集以Apache 2.0许可证发布,适用于法律NLP研究、法规搜索、检索增强生成、文档分类、法律信息提取、法律摘要、法律数据分析和LLM训练与评估等用途。
This dataset, named Indian Legal Documents, is part of the KanoonGPT Open Legal Data Initiative, designed to provide structured Indian statutory and legal-document data for AI, search, and legal research. It contains 35,851 Indian legal documents, including Acts, Rules, Notifications, Circulars, Orders, Regulations, Guidelines, Ordinances, and other legal or quasi-legal government documents, covering jurisdictions such as Central and various states (e.g., Maharashtra, Uttar Pradesh). The time coverage spans from 1970 to 2026, with parsed dates for some documents and others lacking reliably parsed dates. The dataset is stored in an AI-ready, tabular format and supports NLP tasks such as text-generation, question-answering, feature-extraction, text-retrieval, text-ranking, and summarization. The schema includes columns like doc_id, document_title, document_type, document_jurisdiction, issuing_authority, issue_date, short_description, and cleaned text. Released under the Apache License 2.0, it is intended for use in legal NLP research, statutory search, retrieval-augmented generation, document classification, legal information extraction, legal summarization, legal data analytics, and LLM training and evaluation.




