遇见数据集

DrugProt corpus: Biocreative VII Track 1 - Text mining drug and chemical-protein interactions

收藏
Zenodo2023-12-04 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Gold Standard annotations of the DrugProt corpus (training and development sets) <br> <strong>Introduction</strong> The aim of the DrugProt track (similar to the previous CHEMPROT task of BioCreative VI) is to promote the development and evaluation of systems that are able to automatically detect in relations between chemical compounds/drug and genes/proteins. We have therefore generated a manually annotated corpus, the <em>DrugProt corpus</em>, where domain experts have exhaustively labeled:(a) all chemical and gene mentions, and (b) all binary relationships between them corresponding to a specific set of biologically relevant relation types (<em>DrugProt relation classes</em>). There is also an increasing interested in the integration of chemical and biomedical data understood as curation of relationships between biological and chemical entities from text and storing such information in form of structured annotation databases. Such databases are of key relevance not only for biological but also for pharmacological and clinical research. A range of different types chemical-protein/gene interactions are of key relevance for biology, including metabolic relations (e.g. substrates, products) inhibition, binding or induction associations. The DrugProt track aims to address these needs and to promote the development of systems able to extract chemical-protein interactions that might be of relevance for precision medicine as well as for drug discovery and basic biomedical research. The DrugProt track in BioCreative VII (BC VII) will explore recognition of chemical-protein entity relations from abstracts. Teams participating in this track are provided with: PubMed abstracts Manually annotated chemical compound mentions Manually annotated gene/protein mentions Manually annotated chemical compound-protein relations <strong>Zip structure:</strong> Training set folder with drugprot_training_abstracts.tsv: PubMed records drugprot_training_entities.tsv: manually labeled mention annotations of chemical compounds and genes/proteins drugprot_training_relations.tsv: chemical-­protein relation annotations Development set folder with drugprot_development_abstracts.tsv drugprot_development_entities.tsv drugprot_development_relations.tsv <strong>Data format description</strong> The <strong>input text files</strong> for the DrugProt track will be plain-text, UTF8-encoded PubMed records in a tab-separated format with the following three columns: Article identifier (PMID, PubMed identifier) Title of the article Abstract of the article DrugProt <strong>entity mention annotation files</strong> contain manually labeled mention annotations of chemical compounds and genes/proteins. Such files consist of tab-separated fields containing the following six columns: Article identifier (PMID) Term number (for this record) Type of entity mention (CHEMICAL, GENE-Y, GENE-N) Start character offset of the entity mention End character offset of the entity mention Text string of the entity mention Each line contains one entity, and <em>each entity is uniquely identified by its PMID and the Term Number</em>. Besides, each annotation contains an annotation type, the start-offset -the index of the first character of the annotated span in the text-, the end-offset -the index of the first character after the annotated span- and the text spanned by the annotation. Example DrugProt <em>training</em> entity mention annotations: <pre><code>11808879 T1 GENE-Y 1860 1866 KIR6.2 11808879 T2 GENE-N 1993 2016 glutamate dehydrogenase 11808879 T3 GENE-Y 2242 2253 glucokinase 23017395 T1 CHEMICAL 216 223 HMG-CoA 23017395 T2 CHEMICAL 258 261 EPA</code></pre> Example DrugProt <em>development</em> entity mention annotations (no distinction between GENE-Y and GENE-N): <pre><code>11808879 T1 GENE 1860 1866 KIR6.2 11808879 T2 GENE 1993 2016 glutamate dehydrogenase 11808879 T3 GENE 2242 2253 glucokinase 23017395 T1 CHEMICAL 216 223 HMG-CoA 23017395 T2 CHEMICAL 258 261 EPA</code></pre> <br> DrugProt <strong>relation annotations</strong> will be distributed as a file that contains the detailed chemical-protein relation annotations prepared for the DrugProt track. It consists of tab-separated columns containing: Article identifier (PMID) DrugProt relation Interactor argument 1 (<em>of type CHEMICAL</em>) Interactor argument 2 (<em>of type GENE</em>) Each line contains one relation, and <em>each relation is identified by the PMID, the relation type and the two related entities</em>. In the below example, to find the entities involved in the first relation, you must find the entities with Term Identifier T1 and T52 <em>within the PMID 12488248.</em> Example DrugProt relation annotations: <pre><code>12488248 INHIBITOR Arg1:T1 Arg2:T52 12488248 INHIBITOR Arg1:T2 Arg2:T52 23220562 ACTIVATOR Arg1:T12 Arg2:T42 23220562 ACTIVATOR Arg1:T12 Arg2:T43 23220562 INDIRECT-DOWNREGULATOR Arg1:T1 Arg2:T14</code></pre> Please, cite: @inproceedings{krallinger2017overview, title={Overview of the BioCreative VI chemical-protein interaction Track}, author={Krallinger, Martin and Rabal, Obdulia and Akhondi, Saber A and P{\'e}rez, Mart{\i}n P{\'e}rez and Santamar{\'\i}a, Jes{\'u}s and Rodr{\'\i}guez, Gael P{\'e}rez and others}, booktitle={Proceedings of the sixth BioCreative challenge evaluation workshop}, volume={1}, pages={141--146}, year={2017}} <strong>Summary statistics:</strong> <pre><code> Training set Development set Documents 3500 750 Tokens 1001168 199620 Annotated Entities 89529 18858 Annotated Relations 17288 3765</code></pre> Annotated Entities: <pre><code class="language-html"> Training Entities Development Entities CHEMICAL 46274 9853 GENE-Y [Normalizable] 28421 - GENE-N [Non-Normalizable] 14834 - Gene Total (N+Y) 43255 9005 Total 89529 18858</code></pre> Annotated Relations: <pre><code> Training Relations Development Relations INDIRECT-DOWNREGULATOR 1330 332 INDIRECT-UPREGULATOR 1379 302 DIRECT-REGULATOR 2250 458 ACTIVATOR 1429 246 INHIBITOR 5392 1152 AGONIST 659 131 AGONIST-ACTIVATOR 29 10 AGONIST-INHIBITOR 13 2 ANTAGONIST 972 218 PRODUCT-OF 921 158 SUBSTRATE 2003 495 SUBSTRATE_PRODUCT-OF 25 3 PART-OF 886 258 Total 17288 3765</code></pre> For further information, please visit https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/ or email us at krallinger.martin@gmail.com and antoniomiresc@gmail.com <strong>Related resources:</strong> Web Evaluation library Relation annotation guidelines Gene and protein annotation guidelines Chemicals and drugs annotation guidelines FAQ

DrugProt语料库(训练集与开发集)的金标准标注集<br><strong>任务简介</strong><br>DrugProt任务赛道(与BioCreative VI的CHEMPROT任务一脉相承)旨在推动能够自动识别化合物/药物与基因/蛋白质之间关联的系统的研发与评测。为此,我们构建了人工标注语料库<em>DrugProt语料库(DrugProt corpus)</em>,由领域专家完成全量标注:(a) 所有化合物与基因实体提及,(b) 对应于一系列生物学相关关联类型(<em>DrugProt关联类别(DrugProt relation classes)</em>)的全部实体间二元关联。<br>当前学界对化学与生物医学数据的整合需求日益增长,即从文本中抽取生物与化学实体间的关联并将其存储为结构化标注数据库。此类数据库不仅对生物学研究至关重要,同样在药理学与临床研究中具有核心价值。多种类型的化学物-蛋白质/基因相互作用对生物学研究具有核心意义,包括代谢关联(如底物、产物)、抑制、结合或诱导调控关联等。DrugProt任务赛道旨在满足此类需求,推动能够提取化学物-蛋白质相互作用的系统研发,此类相互作用对精准医学、药物研发以及基础生物医学研究均具有重要价值。<br>BioCreative VII(BC VII)中的DrugProt任务赛道将聚焦于从文献摘要中识别化学物-蛋白质实体关联。参与本赛道的团队将获得以下数据:PubMed文献摘要、人工标注的化合物实体提及、人工标注的基因/蛋白质实体提及、人工标注的化学物-蛋白质关联。<br><strong>压缩包结构:</strong>训练集文件夹包含以下文件:drugprot_training_abstracts.tsv:PubMed文献记录;drugprot_training_entities.tsv:化合物与基因/蛋白质的人工标注实体提及;drugprot_training_relations.tsv:化学物-蛋白质关联标注。开发集文件夹包含以下文件:drugprot_development_abstracts.tsv、drugprot_development_entities.tsv、drugprot_development_relations.tsv<br><strong>数据格式说明</strong><br>DrugProt任务赛道的<strong>输入文本文件</strong>为UTF8编码的纯文本格式PubMed文献记录,采用制表符分隔格式,包含以下三列:文献标识符(PMID,PubMed唯一标识符)、文献标题、文献摘要。<br><strong>实体提及标注文件</strong>包含化合物与基因/蛋白质的人工标注实体提及。此类文件采用制表符分隔格式,包含以下六列:文献标识符(PMID)、实体编号(本记录内唯一)、实体提及类型(CHEMICAL、GENE-Y、GENE-N)、实体提及的起始字符偏移量、实体提及的结束字符偏移量、实体提及的文本内容。每行对应一个实体,每个实体通过其PMID与实体编号实现唯一标识。此外,每条标注包含标注类型、起始偏移量(文本中标注片段的首个字符的索引)、结束偏移量(标注片段后首个字符的索引)以及标注覆盖的文本内容。<br>DrugProt <em>训练集</em>实体提及标注示例:<pre><code>11808879 T1 GENE-Y 1860 1866 KIR6.2 11808879 T2 GENE-N 1993 2016 glutamate dehydrogenase 11808879 T3 GENE-Y 2242 2253 glucokinase 23017395 T1 CHEMICAL 216 223 HMG-CoA 23017395 T2 CHEMICAL 258 261 EPA</code></pre>DrugProt <em>开发集</em>实体提及标注示例(不区分GENE-Y与GENE-N):<pre><code>11808879 T1 GENE 1860 1866 KIR6.2 11808879 T2 GENE 1993 2016 glutamate dehydrogenase 11808879 T3 GENE 2242 2253 glucokinase 23017395 T1 CHEMICAL 216 223 HMG-CoA 23017395 T2 CHEMICAL 258 261 EPA</code></pre><br>DrugProt <strong>关联标注文件</strong>将以制表符分隔格式提供,包含为DrugProt任务赛道准备的详细化学物-蛋白质关联标注。该文件包含以下列:文献标识符(PMID)、DrugProt关联类型、相互作用对1(类型为CHEMICAL)、相互作用对2(类型为GENE)。每行对应一条关联,每条关联通过其PMID、关联类型与两个关联实体实现唯一标识。在以下示例中,若要查找第一条关联涉及的实体,需在PMID为12488248的文献中找到实体编号为T1与T52的实体。DrugProt关联标注示例:<pre><code>12488248 INHIBITOR Arg1:T1 Arg2:T52 12488248 INHIBITOR Arg1:T2 Arg2:T52 23220562 ACTIVATOR Arg1:T12 Arg2:T42 23220562 ACTIVATOR Arg1:T12 Arg2:T43 23220562 INDIRECT-DOWNREGULATOR Arg1:T1 Arg2:T14</code></pre><br>请引用:<pre>@inproceedings{krallinger2017overview, title={BioCreative VI化学物-蛋白质相互作用任务赛道综述}, author={Krallinger, Martin and Rabal, Obdulia and Akhondi, Saber A and Pérez, Martín Pérez and Santamaría, Jesús and Rodríguez, Gael Pérez and others}, booktitle={第六届BioCreative挑战赛评估研讨会论文集}, volume={1}, pages={141--146}, year={2017}}</pre><br><strong>统计摘要:</strong><pre><code> 训练集 开发集 文献数 3500 750 词元数 1001168 199620 标注实体数 89529 18858 标注关联数 17288 3765</code></pre><br>标注实体统计:<pre><code class="language-html"> 训练集实体 开发集实体 CHEMICAL 46274 9853 GENE-Y [可标准化] 28421 - GENE-N [不可标准化] 14834 - 基因总计(N+Y) 43255 9005 总计 89529 18858</code></pre><br>标注关联统计:<pre><code> 训练集关联 开发集关联 INDIRECT-DOWNREGULATOR(间接下调) 1330 332 INDIRECT-UPREGULATOR(间接上调) 1379 302 DIRECT-REGULATOR(直接调控) 2250 458 ACTIVATOR(激活) 1429 246 INHIBITOR(抑制) 5392 1152 AGONIST(激动剂) 659 131 AGONIST-ACTIVATOR(激动剂-激活) 29 10 AGONIST-INHIBITOR(激动剂-抑制) 13 2 ANTAGONIST(拮抗剂) 972 218 PRODUCT-OF(产物) 921 158 SUBSTRATE(底物) 2003 495 SUBSTRATE_PRODUCT-OF(底物-产物) 25 3 PART-OF(属于) 886 258 总计 17288 3765</code></pre><br>如需获取更多信息,请访问 https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/ 或发送邮件至 krallinger.martin@gmail.com 与 antoniomiresc@gmail.com<br><strong>相关资源:</strong>Web评估库、关联标注指南、基因与蛋白质标注指南、化合物与药物标注指南、常见问题(FAQ)

提供机构:
Zenodo
创建时间:
2021-06-29
二维码
社区交流群
二维码
科研交流群
商业服务