遇见数据集
数据链接:
官方服务:

资源简介:

SureChEMBL is a publicly available large-scale data source of scientificaly annotated patents.The data are extracted from the patent literature according to an automated text and image-mining pipeline on a daily basis. SureChEMBL provides access to a previously unavailable, open and timely set of annotated compound-patent associations, complemented with sophisticated combined structure and keyword-based search capabilities against the compound repository and patent document corpus. More recently SureChEMBL2.0 introduced text annotation for genes/proteins, diseases and mechanisms of action. As of January 2026, the database contains 32 million unique compounds, and over 1 million biomedical annotations (900 thousand genes/proteins and 138 thousand diseases), extracted from 55 million patent documents.

SureChEMBL是一款可公开获取的大规模科学标注专利数据库。其数据每日通过自动化文本与图像挖掘流水线从专利文献中提取得到。SureChEMBL提供了此前尚未公开的、开放且实时的标注化化合物-专利关联集访问渠道,并针对化合物库与专利文档语料库,配备了精密的结构化与关键词联合检索功能。近期推出的SureChEMBL 2.0版本新增了针对基因/蛋白质、疾病以及作用机制的文本标注功能。截至2026年1月,该数据库已从5500万份专利文档中提取得到3200万种独特化合物,以及超过100万条生物医学标注信息(其中包含90万条基因/蛋白质标注与13.8万条疾病标注)。

搜集汇总
数据集介绍
SureChEMBL 数据集图片
背景与挑战
背景概述
SureChEMBL是一个公开的大规模科学注释专利数据集,通过自动化流程从专利文献中提取化合物-专利关联信息,并支持结构和关键词搜索。其2.0版本扩展了生物医学注释,包括基因/蛋白质和疾病,截至2026年1月涵盖3200万种独特化合物和超过100万项注释。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务