PatenTEB
收藏资源简介:
PatenTEB是一个全面的基准数据集,包含了15个任务,涵盖了检索、分类、释义检测和聚类,拥有206万个示例。该数据集采用了领域分层划分、领域特定的硬负样本挖掘以及系统性地涵盖了常规嵌入基准中缺失的非对称片段到文档匹配场景。PatenTEB旨在解决专利文本嵌入中的挑战,包括长文档的依赖处理、非对称匹配场景和跨领域语义理解等问题。
PatenTEB is a comprehensive benchmark dataset encompassing 15 tasks covering retrieval, classification, paraphrase detection, and clustering, with a total of 2.06 million instances. This dataset adopts domain-stratified partitioning, domain-specific hard negative mining, and systematically incorporates asymmetric segment-to-document matching scenarios that are absent from conventional embedding benchmarks. PatenTEB aims to address core challenges in patent text embedding, including handling contextual dependencies in long documents, asymmetric matching scenarios, and cross-domain semantic understanding.
PatenTEB数据集概述
数据集基本信息
- 名称: PatenTEB (Patent Text Embedding Benchmark)
- 类型: 专利文本嵌入基准数据集
- 许可证: CC BY-NC-SA 4.0
- 数据来源: Lens.org
数据集规模
- 测试集: 319,320个样本(已发布)
- 训练集: 1,556,751个样本(计划发布)
- 验证集: 181,215个样本(计划发布)
- 总计: 2,057,286个样本
任务分类
分类任务(3个)
class_bloom: 引用时序分类class_nli_oldnew: 引用方向性分类class_text2ipc3: IPC3技术分类
聚类任务(2个)
clusters_ext_full_ipc: 基于IPC的聚类clusters_inventor: 基于发明人的聚类
对称检索任务(3个)
retrieval_IN: 同领域检索(相同IPC3)retrieval_MIXED: 混合领域检索(部分IPC3重叠)retrieval_OUT: 跨领域检索(不相交IPC3)
非对称检索任务(5个)
title2full: 标题→全文problem2full: 问题→全文problem2solution: 问题→解决方案effect2full: 效果→全文effect2substance: 效果→实质内容
复述任务(2个)
para_problem: 问题复述检测para_solution: 解决方案复述检测
评估指标
- 分类任务: Macro-F1
- 聚类任务: V-measure
- 检索任务: NDCG@10
- 复述任务: Pearson r
模型性能
- 总体得分: 0.654(PatenTEB基准)
- BigPatentClustering.v2: 0.494 V-measure(新SOTA)
- DAPFAM跨领域专利检索: 0.377 NDCG@100
访问方式
- HuggingFace数据集: https://huggingface.co/datalyes
- 论文: https://arxiv.org/abs/2510.22264
- 模型: https://huggingface.co/datalyes

- 1PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding法国斯特拉斯堡国立应用科学学院(INSA Strasbourg, France) · 2025年



