CASESUMM
收藏资源简介:
CASESUMM是由芝加哥大学研究团队创建的一个大规模法律领域长文本摘要数据集,包含了25,600条美国最高法院的意见书及其官方摘要(syllabuses),时间跨度为1815年至2019年。该数据集是目前最大的公开法律案例摘要数据集,涵盖了超过200年的最高法院判决。数据集的内容包括每个案件的事实、程序历史、法律问题及其解答,摘要由法院雇佣的律师撰写并经法官批准,具有较高的权威性。数据集的创建过程涉及从多个来源(如Public Resource Org和国会图书馆)提取和清理意见书及摘要,并通过OCR技术和正则表达式进行结构化处理。该数据集主要用于评估大语言模型在法律领域的摘要生成能力,旨在解决长文本摘要任务中的复杂性和高要求问题。
CASESUMM is a large-scale legal long-text summarization dataset constructed by a research team at the University of Chicago. It contains 25,600 U.S. Supreme Court opinions and their official syllabuses, spanning the period from 1815 to 2019. This dataset is the largest publicly available legal case summarization dataset to date, covering Supreme Court decisions spanning over 200 years. The dataset includes the factual background, procedural history, legal issues and their resolutions for each case. These official syllabuses are drafted by attorneys employed by the Court and approved by justices, rendering them highly authoritative. The construction of this dataset involved extracting and curating court opinions and syllabuses from multiple sources, including Public Resource Org and the Library of Congress, followed by structured processing via OCR technology and regular expressions. This dataset is primarily utilized to evaluate the summarization performance of large language models (LLMs) in the legal domain, with the goal of addressing the complexity and stringent requirements associated with long-text summarization tasks.




