pkheria7/indian-legal-opposing-counsel-dataset
收藏资源简介:
Indian Legal Opposing Counsel Dataset是一个专门用于训练印度法律对抗律师AI模型的预处理数据集,包含26,326个示例,采用ChatML格式,适用于监督微调(SFT)训练。该数据集整合了三个来源:viber1/indian-law-dataset(24,607行,内容涉及令状、公益诉讼、民事诉讼程序、宪法法律和印度刑法典)、nisaar/Lawyer_GPT_India(150行,内容涵盖里程碑案例、印度刑法典、合同法和宪法原则)以及RMani1/indian-legal-dataset-indian-law(1,569行,内容为印度法规、法案和法律条款)。数据集分为训练集(25,009行,65 MB)和测试集(1,317行,3.5 MB),总大小69 MB。每个数据行包含一个messages列,采用ChatML对话格式,包括系统角色、用户角色和助手角色,模拟法律咨询对话。数据集旨在帮助开发AI模型,使其能够作为印度法律领域的对抗律师,处理宪法、法律程序和相关法律问题。数据集使用Apache 2.0许可证,并提供多种下载方式,如Python加载、直接下载JSONL文件、wget/curl和Git克隆。
Indian Legal Opposing Counsel Dataset is a combined, preprocessed dataset of 26,326 examples for training an Indian legal opposing counsel AI model, ready-to-use in ChatML format for SFT training. It integrates data from three sources: viber1/indian-law-dataset (24,607 rows, covering writs, PIL, civil procedure, constitutional law, IPC), nisaar/Lawyer_GPT_India (150 rows, covering landmark cases, IPC, contract law, constitutional principles), and RMani1/indian-legal-dataset-indian-law (1,569 rows, covering Indian statutes, acts, legal provisions). The dataset is split into train (25,009 rows, 65 MB) and test (1,317 rows, 3.5 MB) sets, totaling 69 MB. Each row features a messages column in ChatML conversational format, with system, user, and assistant roles to simulate legal consultation dialogues. It is designed to develop AI models that can act as opposing counsel in Indian legal contexts, addressing constitutional, procedural, and legal issues. The dataset is licensed under Apache 2.0 and offers multiple download options, including Python loading, direct JSONL downloads, wget/curl, and Git cloning.



