Quantifying Legal Risk in Corporate Speech — Outcome-Linked Quotes Dataset
收藏资源简介:
Quantifying Legal Risk in Corporate Speech — Outcome-Linked Quotes Dataset This dataset contains curated JSONL entries of corporate speech extracted from major U.S. litigation and regulatory proceedings, each linked to subsequent legal outcomes. Records include provenance fields (CourtListener docket/document IDs, source URLs, and jurisdiction metadata), short text quotations, labels, engineered features, and model predictions. Records: 24,462 quotes Cases: 125 cases Time span: 2000–2025 Jurisdictions: 10+ U.S. states Outcomes: Ranging from \$32k settlements to \$5B+ judgments File size: ~1 GB compressed, 3 GB Fields in each row: Embeddings: Large float array (CORAL embeddings). CORAL predictions: bucket, class, confidence, probabilities, scores. Features: 200+ interpretable and engineered features (lexical, sequential, linguistic, structural, legal risk). Metadata: case_id_clean (e.g., "1:19-cv-02184_dcd"), _metadata_src_path, docket_id, courtlistener_id, source_url, jurisdiction. Leakage prevention wrappers: separate quote, speaker, and context fields. Source: Derived from CourtListener (Free Law Project) via the RECAP Archive and API. Only metadata and short quotations necessary for analysis are included. Full filings are not redistributed; entries link back to CourtListener. Use case: Enables research on NLP methods for legal risk quantification, outcome prediction, and explainable AI in corporate compliance, litigation discovery, and insurance premium modeling. Citation: Jacob Dugan (2025). Quantifying Legal Risk in Corporate Speech — Outcome-Linked Quotes Dataset. Zenodo. https://doi.org/10.5281/zenodo.16934610 Attribution: CourtListener (Free Law Project). https://www.courtlistener.com Citation guide: https://www.courtlistener.com/cite/ License: CC BY 4.0 for annotations and derived labels. Original filings remain under CourtListener terms of service and applicable court rules.
企业言论法律风险量化——关联法律结果引述数据集 本数据集收录经整理的JSONL格式条目,数据源自美国主要诉讼与监管程序中的企业言论,每条条目均关联后续法律判决结果。记录包含溯源字段(CourtListener(Free Law Project)案卷/文档ID、来源URL及管辖元数据)、短文本引述、标签、工程化特征与模型预测结果。 数据统计: - 记录总数:24462条引述 - 案件总数:125起 - 时间跨度:2000年至2025年 - 管辖范围:覆盖美国10余个州 - 判决结果跨度:从3.2万美元的和解金额至50亿美元以上的判决金额 - 文件大小:压缩后约1GB,解压后约3GB 单条数据行包含以下字段: 1. 嵌入向量:采用CORAL嵌入(CORAL embeddings)的大型浮点数组 2. CORAL预测结果:包含分类桶、类别、置信度、概率分布与得分 3. 工程化特征:包含200余个可解释的人工构造特征,涵盖词汇、序列、语言、结构与法律风险维度 4. 元数据:包括清理后的案件ID(格式示例:"1:19-cv-02184_dcd")、_metadata_src_path、案卷ID、CourtListener ID、来源URL与管辖区域 5. 防数据泄露封装字段:包含独立的引述、发言者与上下文字段 数据集来源:本数据集源自CourtListener(Free Law Project),通过RECAP档案与API获取。仅收录分析所需的元数据与短文本引述,未分发完整案卷文件,所有条目均链接至CourtListener原平台。 使用场景:本数据集可支持多项研究方向,包括用于法律风险量化、结果预测的自然语言处理方法,以及企业合规、诉讼证据开示与保险费率建模中的可解释AI研究。 引用信息: Jacob Dugan (2025). 企业言论法律风险量化——关联法律结果引述数据集. Zenodo. https://doi.org/10.5281/zenodo.16934610 归属声明: CourtListener(Free Law Project). https://www.courtlistener.com 引用指南:https://www.courtlistener.com/cite/ 许可协议: - 标注内容及衍生标签采用CC BY 4.0协议 - 原始案卷文件仍受CourtListener服务条款与适用法院规则约束



