CaseHOLD

Opencsg2024-07-17 更新2024-07-22 收录

下载链接：

https://www.opencsg.com/datasets/MagicAI/CaseHOLD

下载链接

链接失效反馈

官方服务：

资源简介：

预训练语料库是通过摄取从1965年至今的整个哈佛法学院案例语料库构建的。这个语料库（37GB）的大小很大，代表了所有联邦和州法院的3,446,187个法律判决，并且比最初用于训练BERT的BookCorpus/Wikipedia语料库（15GB）的大小还要大。我们从这个语料库中随机抽取 10% 的决策作为保留集，我们用它来创建 CaseHOLD 数据集。剩下的 90% 用于预训练。

The pre-training corpus was constructed by ingesting the entire Harvard Law School case corpus from 1965 to the present. This corpus, with a size of 37GB, is substantial, encompassing 3,446,187 legal judgments from all federal and state courts, and is larger than the BookCorpus/Wikipedia corpus (15GB) originally used to train BERT. We randomly sampled 10% of the decisions from this corpus as a held-out set, which we employed to create the CaseHOLD dataset. The remaining 90% was utilized for pre-training.

创建时间：

2024-07-17

搜集汇总

数据集介绍

背景与挑战

背景概述

CaseHOLD数据集包含53,000多个法律决策的多选问答，用于评估预训练模型在法律领域的表现。该数据集基于哈佛法学院案例语料库构建，覆盖了1965年至今的联邦和州法院的3,446,187个法律判决。

以上内容由遇见数据集搜集并总结生成

5,000+

优质数据集

54 个

任务类型

进入经典数据集