ContractScrub
收藏资源简介:
ContractScrub是由汤森路透基础研究与帝国理工学院联合创建的首个专用于评估法律合同终审校对的基准数据集。该数据集包含3,014个标注任务,涵盖44份合同中的九类典型错误,如定义术语误用、引用错误和语言不一致等,均由资深律师手工标注并注入额外错误以构建黄金标准。数据来源为公开的CUAD合约库,经过严格筛选、现有错误标注及针对性错误注入三阶段流程生成。该基准旨在衡量大语言模型在真实合同审查场景中的缺陷识别能力,填补了现有法律推理基准与经济效益之间的空白,为自动化合同校对提供可靠评估。
ContractScrub is the first benchmark dataset specifically designed for evaluating final proofreading of legal contracts, jointly created by Thomson Reuters Basic Research and Imperial College London. This dataset includes 3,014 annotation tasks, covering nine typical error categories across 44 contracts, such as misused defined terms, citation errors, and linguistic inconsistencies. All annotation work was manually completed by senior lawyers, with additional targeted errors injected to establish the gold standard. Derived from the publicly available CUAD contract corpus, the dataset was generated through a three-stage workflow consisting of strict screening, annotation of existing errors, and targeted error injection. This benchmark aims to evaluate the defect detection capability of Large Language Models (LLMs) in real-world contract review scenarios, filling the gap between existing legal reasoning benchmarks and practical economic benefits, and providing a reliable evaluation framework for automated contract proofreading.
数据集概述:tri-fair-lab/contract_scrub
- 数据集名称:contract_scrub
- 提供机构:tri-fair-lab
- 许可证:CC BY-NC 4.0(知识共享-署名-非商业性使用 4.0 国际许可协议)
- 当前状态:该数据集详情页面标注为“即将推出”(Coming soon!),目前尚未正式发布具体的数据内容、格式或使用说明。
- 访问地址:https://huggingface.co/tri-fair-lab/contract_scrub
该页面目前仅包含基础的许可信息与占位说明,无更多关于数据规模、字段定义、采集方式或应用场景的详细描述。





