overthelex/indian-court-decisions
收藏资源简介:
Indian Court Decisions 是一个大规模印度法院判决数据集,包含完整文本、元数据和结果标签,覆盖印度最高法院和25个高等法院的判决(时间范围为1950年至2026年)。该数据集是最大的公开法律NLP数据集之一,拥有超过1460万条法院判决记录,提取了全文内容。它涵盖了75年来的印度法理学,涉及所有主要法院,包括高等法院和最高法院的具体配置(例如,高等法院配置有11,682,776条训练记录,最高法院配置有40,044条训练记录)。数据集支持多种语言,主要为英语,但也包括印地语、泰米尔语、泰卢固语、卡纳达语、马拉地语、孟加拉语、马拉雅拉姆语、古吉拉特语、奥里亚语和旁遮普语等印度语言。字段包括案例编号、法院代码、法院名称、法官、全文文本、文本长度、标题、当事人信息、判决日期、处置性质和标准化结果标签等。结果标签涵盖多种类别,如驳回、处置、允许、部分允许、撤回、关闭、保释批准、保释拒绝、撤销、转移、和解、命令和其他。数据集可用于判决预测、法律文本分类、时间分析、跨司法管辖区比较、法律语言建模和信息检索等应用。数据来源于AWS开放数据注册表,由Dattam Labs提供,原始数据收集自eCourts和印度最高法院。
Indian Court Decisions is a large-scale dataset of Indian court decisions with full text, metadata, and outcome labels covering the Supreme Court of India and 25 High Courts (1950–2026). It is one of the largest publicly available legal NLP datasets, containing over 14.6 million court decisions with extracted full text. The dataset covers 75 years of Indian jurisprudence across all major courts, including specific configurations for High Courts (e.g., 11,682,776 training records) and Supreme Court (e.g., 40,044 training records). It supports multiple languages, primarily English, but also includes Hindi, Tamil, Telugu, Kannada, Marathi, Bengali, Malayalam, Gujarati, Odia, and Punjabi. Fields include case number, court code, court name, judge, full text, text length, title, petitioner, respondent, decision date, disposal nature, and normalized outcome labels. Outcome labels cover categories such as dismissed, disposed, allowed, partly allowed, withdrawn, closed, bail granted, bail rejected, quashed, transferred, settled, ordered, and other. The dataset is useful for judgment prediction, legal text classification, temporal analysis, cross-jurisdictional comparison, legal language modeling, and information retrieval. Data is sourced from the AWS Open Data Registry by Dattam Labs, with original data collected from eCourts and the Supreme Court of India.




