DrugProt Complete PubMed Knowledge Graph
收藏资源简介:
DrugProt Complete PubMed Knowledge Graph Please cite if you use any DrugProt resource: Antonio Miranda-Escalada, Farrokh Mehryary, Jouni Luoma, Darryl Estrada-Zavala, Luis Gasco, Sampo Pyysalo, Alfonso Valencia, Martin Krallinger, Overview of DrugProt task at BioCreative VII: data and methods for large-scale text mining and knowledge graph generation of heterogenous chemical–protein relations, Database, Volume 2023, 2023, baad080 @article{miranda2023overview, title={Overview of DrugProt task at BioCreative VII: data and methods for large-scale text mining and knowledge graph generation of heterogenous chemical--protein relations}, author={Miranda-Escalada, Antonio and Mehryary, Farrokh and Luoma, Jouni and Estrada-Zavala, Darryl and Gasco, Luis and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, journal={Database}, volume={2023}, pages={baad080}, year={2023}, publisher={Oxford University Press UK} } Description This dataset contains a knowledge graph built from PubMed dump abstracts (December 2021). A NER system has been applied to each of the abstracts to extract mentions of type "CHEMICAL" and "GENE", as well as a RE system to detect existing relations between these mentions such as ACTIVATOR, INHIBITOR, AGONIST or PRODUCT_OF, among others (see article for a full list of relations considered). Given the volume of the dataset, the repository is divided into 1114 folders. Each of these folders contains a chunk of PubMed abstracts, entities and relationships, divided into the following 3 files: abstracts.tsv: Tabular file in which each line represents a pubmed document. The file has 3 columns: Pubmed_id: Numerical identifier of the document in PubMed Title: Title of the document Abstract: Abstract text. entities.tsv: List of the entities extracted from the abstracts. Each line represents an extracted entity, and has 5 columns: Pubmed_id: Numerical identifier of the document in PubMed Mention_id: Numerical identifier of the mention in the document. Entity_type: Type of mention. It can be CHEMICAL or GENE. Span_ini: Index of the first character of the annotated span in the text Span_end: Index of the first character after the annotated span. Span: Text span of the annotation relations.tsv: File of existing relations between entities. Each line represents a relationship, and has the following fields: Pubmed_id: Numerical identifier of the document in PubMed Relation_type: DrugProt relation type among entities/arguments. Arg1: Mention of CHEMICAL Arg2: Mention of GENE Files: drugprot-silver-standard-kg.zip : Folders with the files previously explained Related resources: Web DrugProt corpus Evaluation library Online evaluation (CodaLab) Relation annotation guidelines Gene and protein annotation guidelines Chemicals and drugs annotation guidelines FAQ DrugProt Large Scale Additional SubTrack DrugProt Large Scale document collection protocol DrugProt Complete PubMed Knowledge Graph



