剑桥法律语料库 (CLC)
收藏资源简介:
剑桥法律语料库(CLC)是由剑桥大学创建的一个大型法律AI研究数据集,包含超过25万个英国法院案例,覆盖时间从16世纪至21世纪。数据集主要由英国法院的判决文本组成,包括案件的基本信息和详细判决内容。创建过程中,原始的Word和PDF文件被转换为XML格式,以便于结构化存储和分析。该数据集特别适用于法律AI研究,如案件结果预测、法律实体识别等,旨在通过机器学习模型自动化解决法律领域的关键任务。
The Cambridge Law Corpus (CLC) is a large-scale legal AI research dataset developed by the University of Cambridge, containing over 250,000 UK court cases spanning from the 16th to the 21st century. The corpus primarily comprises judgment texts from UK courts, including basic case information and detailed adjudication contents. During its development, original Word and PDF files were converted to XML format to facilitate structured storage and analysis. This dataset is particularly suited for legal AI research such as case outcome prediction, legal entity recognition, and other related tasks, aiming to automate key tasks in the legal domain via machine learning models.



