合成法律推理数据集
收藏资源简介:
合成法律推理数据集是由南京大学的研究团队使用KGDG框架生成的,包含5万个高质量的法律推理任务示例。数据集基于一个包含刑事和民事法律文书的知识库构建,通过引导生成具有问题-答案对和推理路径的合成数据,并经过验证和修正以确保质量。该数据集旨在提升开源LLM模型在法律推理任务上的性能,并已公开提供以促进未来研究。
The Synthetic Legal Reasoning Dataset, generated by a research team from Nanjing University using the KGDG framework, consists of 50,000 high-quality legal reasoning task examples. Built upon a knowledge base containing criminal and civil legal documents, this dataset produces synthetic data with question-answer pairs and reasoning paths via guided generation, and has been validated and revised to guarantee its quality. This dataset aims to improve the performance of open-source large language models (LLMs) on legal reasoning tasks, and has been publicly made available to promote future research.

- 1LawGPT: Knowledge-Guided Data Generation and Its Application to Legal LLM南京大学 · 2025年



