Dis2Pat
收藏资源简介:
Dis2Pat是由剑桥大学等机构构建的披露到专利数据集,旨在模拟真实专利起草流程,输入为去法律化的发明者风格披露及附图,输出完整专利申请(权利要求和说明书)。数据集包含约9,433条样本,平均原始专利含11,207个tokens而伪披露仅1,196个tokens,通过从PatentDesc筛选授权专利并利用GPT-5-mini生成伪披露创建,经人工评估在幻觉、细节缺失、矛盾及去法律化方面均获高分(>9.7)。该数据集用于评估大型语言模型在长文本、法律约束下的专利生成能力,解决从非结构化披露生成合规专利的核心挑战。
Dis2Pat is a disclosure-to-patent dataset constructed by institutions including the University of Cambridge. It aims to simulate the real patent drafting workflow, taking de-legalized inventor-style disclosures and accompanying drawings as input, and outputting complete patent applications including claims and specifications. The dataset contains approximately 9,433 samples, with an average of 11,207 tokens per original patent and only 1,196 tokens per pseudo-disclosure. It was developed by screening granted patents from the PatentDesc dataset and generating pseudo-disclosures using GPT-5-mini. Human evaluation shows that it achieved high scores (>9.7) across four aspects: hallucination, missing details, inconsistencies, and de-legalization. This dataset is used to evaluate the patent generation capabilities of Large Language Models (LLMs) under long-text and legally constrained scenarios, addressing the core challenge of generating compliant patent applications from unstructured disclosures.
Patent-MAF 数据集详情
基本信息
- 数据集名称: Patent-MAF
- 数据集地址: https://github.com/scylj1/Patent-MAF
内容概述
- 该数据集涉及专利领域,具体为 Patent-MAF(专利MAF)。
- 当前页面提供数据和代码,但状态显示为“即将发布”(coming soon)。
当前可用性
- 代码与数据尚未公开,处于筹备阶段,暂无法获取或使用具体内容。




