DISC-Law-SFT
收藏资源简介:
DISC-Law-SFT数据集专注于中文法律智能系统,旨在提升法律文本理解和生成能力。它包含DISC-Law-SFT-Pair和DISC-Law-SFT-Triplet两个子集,总规模约为40万条数据,主要覆盖法律信息抽取、法律判决预测、法律文档摘要和法律问答等多种法律场景。数据来源于法律专业领域,并经过了标注处理,可用于监督微调任务,以增强模型对外部法律知识的利用能力。该数据集采用Apache-2.0授权许可。
The DISC-Law-SFT dataset focuses on Chinese legal intelligent systems, aiming to enhance the capabilities of legal text understanding and generation. It contains two subsets, DISC-Law-SFT-Pair and DISC-Law-SFT-Triplet, with a total scale of approximately 400,000 data instances. It mainly covers various legal scenarios including legal information extraction, legal judgment prediction, legal document summarization, and legal question answering. The data is sourced from the legal professional domain and has been annotated, which can be used for supervised fine-tuning tasks to enhance the model's ability to leverage external legal knowledge. This dataset is licensed under Apache-2.0.




