DISC-Law-SFT
收藏资源简介:
DISC-Law-SFT数据集专注于中文法律智能系统,旨在提升法律文本理解和生成能力。它包含DISC-Law-SFT-Pair和DISC-Law-SFT-Triplet两个子集,总规模约为40万条数据,主要覆盖法律信息抽取、法律判决预测、法律文档摘要和法律问答等多种法律场景。数据来源于法律专业领域,并经过了标注处理,可用于监督微调任务,以增强模型对外部法律知识的利用能力。该数据集采用Apache-2.0授权许可。
DISC-Law-SFT dataset is dedicated to Chinese legal intelligent systems, aiming to enhance the capabilities of legal text understanding and generation. It comprises two subsets, namely DISC-Law-SFT-Pair and DISC-Law-SFT-Triplet, with a total scale of approximately 400,000 data entries. It mainly covers various legal scenarios including legal information extraction, legal judgment prediction, legal document summarization, and legal question answering. The data is sourced from the professional legal field and has undergone annotation processing, which can be used for supervised fine-tuning tasks to boost the model's ability to leverage external legal knowledge. This dataset is licensed under Apache-2.0.



