NLP4Pheno v1.0: Bacterial Phenotype Predictions Dataset
收藏资源简介:
This dataset contains the complete output from the NLP4Pheno pipeline, including Named Entity Recognition (NER) predictions, Relation Extraction (RE) predictions, and XGBoost-based phenotype predictions for 121,042 bacterial strains extracted from PubMed Central literature. The dataset comprises: - Named Entity Recognition predictions for 9 entity types (STRAIN, SPECIES, ISOLATE, COMPOUND, MEDIUM, ORGANISM, PHENOTYPE, EFFECT, DISEASE) - Relation Extraction predictions for 16 relationship types (primarily STRAIN-centered) - 1,046,299 aggregated strain-phenotype-compound relationships - XGBoost predictions integrating genomic features (protein domains) with text-mined relationships - Network representations suitable for graph analysis This dataset corresponds to the manuscript: "Integrating natural language processing and genome analysis enables accurate bacterial phenotype prediction"
本数据集包含NLP4Pheno流水线的完整输出结果,涵盖命名实体识别(Named Entity Recognition, NER)预测结果、关系抽取(Relation Extraction, RE)预测结果,以及针对从PubMed Central(PubMed中央文献库)文献中提取的121042株细菌的、基于XGBoost的表型预测结果。 该数据集包含以下内容: - 覆盖9种实体类型(STRAIN、SPECIES、ISOLATE、COMPOUND、MEDIUM、ORGANISM、PHENOTYPE、EFFECT、DISEASE)的命名实体识别预测结果 - 针对16种关系类型(以菌株为核心的关系占绝大多数)的关系抽取预测结果 - 1046299条经聚合的菌株-表型-化合物关系数据 - 整合基因组特征(蛋白质结构域)与文本挖掘所得关系的XGBoost预测结果 - 适用于图分析的网络表征形式 本数据集对应研究论文:《整合自然语言处理与基因组分析实现精准细菌表型预测》(Integrating natural language processing and genome analysis enables accurate bacterial phenotype prediction)



