遇见数据集

CyCrawwler/LegalBrain-Indic-Legal-Corpus

收藏
Hugging Face2026-05-05 更新2026-05-31 收录
官方服务:

资源简介:

LegalBrain Indic Legal Corpus 是一个大规模多语言印度法律数据集,旨在支持领域特定的大语言模型训练、法律问答、政策推理与案例检索以及法律工作流自动化的代理系统研究。该数据集包含从多个印度语言的公开法律来源提取的文本,包括英语、印地语、马拉地语、孟加拉语、坎纳达语、泰米尔语、泰卢固语、奥里亚语等。数据集经过预处理和监督对齐,以(context, question, response)三元组格式提供,适用于监督微调、检索增强生成管道和对话式法律助手训练。数据来源仅限于公开可访问的法律资源,如最高法院判决、高等法院决定、法律委员会报告、公共法律教科书和评论、开放法律新闻档案、公共领域法律问答门户以及政府法案、规则和通知。数据集通过清理和标准化流程处理,包括HTML和样板文本去除、OCR和文本校正、语言检测与分割以及去重,并使用Argilla平台进行人工反馈和模型辅助注释,以确保数据质量,训练模型时能准确引用法律内容而非产生幻觉。

LegalBrain Indic Legal Corpus is a large-scale multilingual Indian legal dataset curated to support research in domain-specific LLM training, legal question answering, policy reasoning & case retrieval, and agentic systems for legal workflow automation. This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others. The corpus is structured and processed to be directly usable for supervised fine-tuning, RAG pipelines, and conversational legal assistants, provided in the format of (context, question, response) triplets. Data was collected only from publicly and legally accessible sources, such as Supreme Court judgments, High Court decisions, law Commission reports, public legal textbooks & commentaries, open legal news archives, public domain legal Q&A portals, and government acts, rules, and notifications. The dataset underwent a cleaning and normalization pipeline, including HTML and boilerplate removal, OCR and text correction, language detection and segmentation, and de-duplication, and was constructed using Argilla for human feedback and model-assisted annotation to ensure responses are grounded in context and avoid hallucinations.

提供机构:
CyCrawwler
二维码
社区交流群
二维码
科研交流群
商业服务