NLP4Pheno v1.0: Bacterial Phenotype Predictions Dataset
收藏资源简介:
This dataset contains the complete output from the NLP4Pheno pipeline, including Named Entity Recognition (NER) predictions, Relation Extraction (RE) predictions, and XGBoost-based phenotype predictions for 121,042 bacterial strains extracted from PubMed Central literature. The dataset comprises: - Named Entity Recognition predictions for 9 entity types (STRAIN, SPECIES, ISOLATE, COMPOUND, MEDIUM, ORGANISM, PHENOTYPE, EFFECT, DISEASE) - Relation Extraction predictions for 16 relationship types (primarily STRAIN-centered) - 1,046,299 aggregated strain-phenotype-compound relationships - XGBoost predictions integrating genomic features (protein domains) with text-mined relationships - Network representations suitable for graph analysis This dataset corresponds to the manuscript: "Integrating natural language processing and genome analysis enables accurate bacterial phenotype prediction"



