theatticusproject/cuad-qa
收藏资源简介:
CUAD(Contract Understanding Atticus Dataset)是一个专门用于法律合同审查的自然语言处理数据集,包含510份商业法律合同中的超过13,000个标签,涵盖了41个重要条款类别。该数据集由专家生成,主要用于支持法律合同审查的NLP研究和开发。数据集的创建过程包括法律学生的培训、手动标签、关键词搜索、类别报告审查、律师审查等多个步骤,以确保注释的准确性。数据集仅包含英文样本,且已分为训练集和测试集。
CUAD (Contract Understanding Atticus Dataset) is a natural language processing (NLP) dataset dedicated to legal contract review. It contains over 13,000 annotated labels across 510 commercial legal contracts, covering 41 critical clause categories. This dataset was developed by domain experts, and is primarily designed to support NLP research and development for legal contract review. The dataset construction process includes multiple steps to ensure annotation accuracy, such as training law students, manual labeling, keyword search, category report review, and lawyer review. The dataset exclusively contains English-language samples, and has been split into training and test sets.
数据集概述
名称: CUAD (Contract Understanding Atticus Dataset)
语言: 英语
许可证: CC-BY-4.0
多语言性: 单语种
大小: 10K<n<100K
源数据集: 原始数据
任务类别: 问答
任务ID:
- closed-domain-qa
- extractive-qa
训练与评估索引:
- 配置: default
- 任务: question-answering
- 任务ID: extractive_question_answering
- 分割:
- 训练分割: train
- 评估分割: test
- 列映射:
- 问题: question
- 上下文: context
- 答案:
- 文本: text
- 答案开始位置: answer_start
- 指标:
- 类型: cuad
- 名称: CUAD
数据集结构
特征:
- id: 字符串类型
- title: 字符串类型
- context: 字符串类型
- question: 字符串类型
- answers: 序列类型,包含:
- text: 字符串类型
- answer_start: int32类型
分割:
- 训练集: 22450个样本
- 测试集: 4182个样本
数据集创建
源数据:
- 包含510份商业合同,来自25种不同类型的合同。
注释:
- 由法律学生和律师进行多步骤注释过程,确保准确性。
个人和敏感信息:
- 部分合同条款因保护机密性而被编辑。
数据集使用考虑
社会影响: 未提供详细信息
偏见讨论: 未提供详细信息
其他已知限制: 未提供详细信息




