UniPS
收藏资源简介:
Chinese-Legal-Instruction是一个高质量的中文法律指令数据集,专为训练和评估大型语言模型(LLM)的法律指令遵循能力而构建。它提供丰富、多样且贴近实际应用的中文法律指令数据,支持法律领域自然语言处理任务的研究与开发。数据来源包括从现有法律数据集中筛选重构的样本和专业人员基于真实法律案例、法规条文编写的高质量指令-输出对,采用CC0 1.0公共领域贡献许可证发布。数据集包含三个核心部分:legal_instruction_data(约10,000个样本,采用标准指令遵循格式,包括instruction、input和output字段)、legal_corpus(法律法规、司法文书等文本语料库)和legal_qa_data(法律问答数据集)。数据覆盖民法、刑法、行政法、合同法、知识产权法等多个法律领域,指令类型涵盖法律信息检索、文本摘要、咨询问答、文书起草、案例分析和法律推理等任务,难度从简单查询到复杂推理。适用于法律文本生成、问答、检索分类、推理等NLP任务,可通过HuggingFace datasets库加载,格式为JSONL。注意:数据集仅供学术研究、技术验证和模型训练使用,不构成正式法律意见,使用者应结合专业法律知识并遵守法律法规。
Chinese-Legal-Instruction is a high-quality Chinese legal instruction dataset, specifically designed for training and evaluating the legal instruction-following capabilities of large language models (LLMs). It provides rich, diverse, and practically relevant Chinese legal instruction data to support research and development in legal natural language processing tasks. Data sources include carefully selected and reformatted samples from existing legal datasets, as well as high-quality instruction-output pairs manually crafted by professionals based on real legal cases, regulations, and scenarios, released under the CC0 1.0 Public Domain Dedication license. The dataset consists of three core components: legal_instruction_data (approximately 10,000 samples in standard instruction-following format with instruction, input, and output fields), legal_corpus (a corpus of legal-related texts such as laws, judicial documents, and legal papers), and legal_qa_data (a legal question-answering dataset). It covers multiple legal domains including civil law, criminal law, administrative law, contract law, intellectual property law, etc., with diverse instruction types such as legal information retrieval, text summarization, consultation Q&A, document drafting, case analysis, and legal reasoning, spanning difficulty levels from simple factual queries to complex multi-step reasoning. Suitable for various legal NLP tasks like text generation, Q&A, retrieval and classification, and reasoning, it can be loaded via the HuggingFace datasets library in JSONL format. Note: The dataset is intended for academic research, technical validation, and model training only, does not constitute formal legal advice, and users should combine it with professional legal knowledge and comply with local laws and regulations.
由于该数据集详情页的README内容仅包含许可证信息(cc0-1.0),未提供数据集的名称、描述、用途、构成等具体内容,因此无法生成有效的数据集概述。




