Deciphering sepsis molecular subtypes using large-scale data to identify subtype-specific drug repurposing
收藏资源简介:
Rationale: Sepsis is a life-threatening syndrome that can quickly cause organ failure and death if untreated. Patient variability complicates therapy development, and its molecular mechanisms remain poorly understood. Advancing knowledge of these mechanisms is essential for effective treatments and clearer phenotype definitions. Objectives: To address this issue, we created a large-scale transcriptomic atlas of publicly available adult sepsis patient data with which we performed molecular phenotyping to evaluate patterns of gene expression and identified potential phenotype-specific drug therapies. Methods: We identified studies of bacterial sepsis in three genomic databases (SRA, GEO, and refine.bio) and reviewed metadata from each study to ensure samples met our predetermined inclusion criteria. We combined this data into one large comprehensive data atlas and harmonized gene expression and associated metadata from each sample. Molecular phenotypes of sepsis were identified via clustering analysis of this sepsis transcriptomic data atlas. We then examined clinical correlates of each phenotype and identified gene signatures associated with each. We performed gene set enrichment analysis on those signatures and identified phenotype-specific potential drug repurposing candidates. We also evaluated the associations between computed phenotypes and mortality. Measurements and Main Results: We harmonized data from 3,713 samples across 28 data sets, of which 2,251 were sepsis patients. Clustering analysis identified four phenotypes within the data. We identified statistically significant phenotype associations with survival, disease, and age. Pathway analysis revealed that MHC class II functions, DNA damage, homeostatic pathways and coagulation may characterize underlying response phenotypes, and may have the potential to guide drug development of sepsis therapeutics. Conclusions: We created the largest transcriptomic sepsis atlas to date, from which we identified four molecular sepsis phenotypes. We described underlying dysregulated molecular mechanisms of these phenotypes, associated clinical covariates, and several potential candidate therapies specific to each phenotype. Future studies should seek to validate such drug-phenotype links to advance sepsis precision medicine. Overall design: Meta-analysis of sepsis transcriptomic studies for adult sepsis patients with bulk transcriptomic data. The datasets used for analysis in this study consists of 1 new dataset of RNA-Seq sepsis samples as well as samples from 27 reanalyzed studies from one of three data repositories: GEO, SRA, or refine.bio. The pipeline for preprocessing of reanalyzed samples differs slightly depending on the source repository. Processing information for each study is described below. All samples were adult patients who met sepsis/septic shock criteria and had bulk RNA-Seq expression data from whole blood with blood being drawn within 48 hours. Processing of reanalyzed studies: Reanalyzed studies (10 total) from GEO: GSE100159 - All samples available in the expression matrix for GSE100159 on GEO were reanalyzed. The expression matrix and associated platform mapping probe IDs were downloaded from GEO. Probe IDs were mapped to associated value in expression matrix in R. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE13015 - All samples available in the expression matrix for GSE13015 on GEO were reanalyzed. The expression matrix and associated platform mapping probe IDs were downloaded from GEO. Probe IDs were mapped to associated value in expression matrix in R. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE131761 - All samples available in the expression matrix for GSE131761 on GEO were reanalyzed. The expression matrix and associated platform mapping probe IDs were downloaded from GEO. Probe IDs were mapped to associated value in expression matrix in R. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE137340 - All samples available in the expression matrix for GSE137340 on GEO were reanalyzed. The expression matrix and associated platform mapping probe IDs were downloaded from GEO. Probe IDs were mapped to associated value in expression matrix in R. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE236713 - All samples available in the expression matrix for GSE236713 on GEO were reanalyzed. The expression matrix and associated platform mapping probe IDs were downloaded from GEO. Probe IDs were mapped to associated value in expression matrix in R. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE236892 - All samples available in the expression matrix for GSE236892 on GEO were reanalyzed. The expression matrix was downloaded from GEO. Ensembl gene IDs were mapped to HUGO symbols in R using biomaRt (v2.56.1) with ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE32707 - All samples available in the expression matrix for GSE32707 on GEO were reanalyzed. The expression matrix and associated platform mapping probe IDs were downloaded from GEO. Probe IDs were mapped to associated value in expression matrix in R. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE69063 - All samples available in the expression matrix for GSE69063 on GEO were reanalyzed. The expression matrix and associated platform mapping probe IDs were downloaded from GEO. Probe IDs were mapped to associated value in expression matrix in R. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE134347 - All samples available in the expression matrix for GSE134347 on GEO were reanalyzed. The expression matrix and associated platform mapping probe IDs were downloaded from GEO. Probe IDs were mapped to associated value in expression matrix in R. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE189400 - All samples available in the expression matrix for GSE189400 on GEO were reanalyzed. The expression matrix was downloaded from GEO. Ensembl gene IDs were mapped to HUGO symbols in R using biomaRt (v2.56.1) with ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. Reanalyzed studies (11 total) from SRA: SRP198776 - All RNA-Seq reads available on SRA for SRP198776 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE154918 - All RNA-Seq reads available on SRA for GSE154918 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE185263 - All RNA-Seq reads available on SRA for GSE185263 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE196117 - All RNA-Seq reads available on SRA for GSE196117 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE199816 - All RNA-Seq reads available on SRA for GSE199816 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE211210 - All RNA-Seq reads available on SRA for GSE211210 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE216902 - All RNA-Seq reads available on SRA for GSE216902 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE222393 - All RNA-Seq reads available on SRA for GSE222393 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE232404 - All RNA-Seq reads available on SRA for GSE232404 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE232753 - All RNA-Seq reads available on SRA for GSE232753 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE63311 - All RNA-Seq reads available on SRA for GSE63311 were reanalyzed. The following processing pipeline used for available reads followed the pipeline defined by refine.bio. Reads for each available sample in the dataset were initially downloaded from SRA using prefetch and fasterq-dump. Fastp was used for read trimming and quality control. Transcript quantification was performed using Salmon. The transcript level estimates for the expression matrix were done using tximport in R. Gene ananotation was done with Ensembl gene annotations. Duplicate gene HUGO symbol IDs were merged by taking the mean of expression as is done in refine.bio processing. Samples were then quantile normalized in R using preprocess Core (v1.62.1) with the reference distribution provided by refine.bio (downloaded from https://api.refine.bio/v1/qn_targets/ on June 28th 2022). This expression matrix was combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. Reanalyzed studies (6 total) from refine.bio: GSE33118 - All samples available in the expression matrix for GSE33118 on refine.bio were reanalyzed. The expression matrix was downloaded from refine.bio, no transformations were selected/performed at the time of download. Quantile normalization is performed as part of the refine.bio download process. The dataset was downloaded from refine.bio independently (not as part of a larger downloaded dataset) in order to ensure quantile normalization is performed independently on the dataset. After download Ensembl gene names were mapped to HUGO Symbols using biomaRt (v2.56.1) with ensembl gene annotations. This expression matrix was then combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE57065 - All samples available in the expression matrix for GSE57065 on refine.bio were reanalyzed. The expression matrix was downloaded from refine.bio, no transformations were selected/performed at the time of download. Quantile normalization is performed as part of the refine.bio download process. The dataset was downloaded from refine.bio independently (not as part of a larger downloaded dataset) in order to ensure quantile normalization is performed independently on the dataset. After download Ensembl gene names were mapped to HUGO Symbols using biomaRt (v2.56.1) with ensembl gene annotations. This expression matrix was then combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE65682 - All samples available in the expression matrix for GSE65682 on refine.bio were reanalyzed. The expression matrix was downloaded from refine.bio, no transformations were selected/performed at the time of download. Quantile normalization is performed as part of the refine.bio download process. The dataset was downloaded from refine.bio independently (not as part of a larger downloaded dataset) in order to ensure quantile normalization is performed independently on the dataset. After download Ensembl gene names were mapped to HUGO Symbols using biomaRt (v2.56.1) with ensembl gene annotations. This expression matrix was then combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE66890 - All samples available in the expression matrix for GSE66890 on refine.bio were reanalyzed. The expression matrix was downloaded from refine.bio, no transformations were selected/performed at the time of download. Quantile normalization is performed as part of the refine.bio download process. The dataset was downloaded from refine.bio independently (not as part of a larger downloaded dataset) in order to ensure quantile normalization is performed independently on the dataset. After download Ensembl gene names were mapped to HUGO Symbols using biomaRt (v2.56.1) with ensembl gene annotations. This expression matrix was then combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE74224 - All samples available in the expression matrix for GSE74224 on refine.bio were reanalyzed. The expression matrix was downloaded from refine.bio, no transformations were selected/performed at the time of download. Quantile normalization is performed as part of the refine.bio download process. The dataset was downloaded from refine.bio independently (not as part of a larger downloaded dataset) in order to ensure quantile normalization is performed independently on the dataset. After download Ensembl gene names were mapped to HUGO Symbols using biomaRt (v2.56.1) with ensembl gene annotations. This expression matrix was then combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis. GSE95233 - All samples available in the expression matrix for GSE95233 on refine.bio were reanalyzed. The expression matrix was downloaded from refine.bio, no transformations were selected/performed at the time of download. Quantile normalization is performed as part of the refine.bio download process. The dataset was downloaded from refine.bio independently (not as part of a larger downloaded dataset) in order to ensure quantile normalization is performed independently on the dataset. After download Ensembl gene names were mapped to HUGO Symbols using biomaRt (v2.56.1) with ensembl gene annotations. This expression matrix was then combined with all other datasets included in the study and batch corrected using ComBat from the sva (v3.48) package for final use in analysis.



