Geneformer-guided multi-omics integration identifies Pbx1 as a network hub of hematopoietic stem cell aging
收藏资源简介:
Hematopoietic stem cells (HSCs) constitute an organized hematopoietic system that undergoes age-related alterations, including increased platelet production and decreased erythropoiesis. The fundamental mechanisms driving these shifts remain incompletely understood. We used single-cell RNA-sequencing data to show that old HSCs contain two distinct transcriptional programs: one shared with megakaryocytes and the other reflecting the most primitive HSC state. Developmental time-series profiling further suggests that the acquisition of these programs begins early in life, with the primitive module rising prenatally and megakaryocytic priming emerging after birth. Using a fine-tuned Geneformer (transformer-based deep learning model) to capture higher-order differences between young and old HSCs, coupled with transcriptomic and epigenetic profiling, and transcription factor screens, we identified Pbx1 as a key regulator of these age-related transcriptional and differentiation changes. Specifically, Pbx1 suppresses erythroid differentiation by repressing Gata1 expression. These findings provide insight into HSC aging and may inform approaches to modulate age-associated HSC dysfunction. Repository structure scRNAseq_preprocess/: Preprocessing pipeline for generating loom files from integrated Seurat objects for Geneformer input. scRNAseq_This_study/: Count matrix and analysis pipeline for scRNA-seq data generated in this study (GSE286966; Fig. 1 and fig. S1 and S5H, J and K). scRNAseq_Dataset2/: Analysis pipeline for merging GSE59114 (C57BL/6), GSE59114 (DBA), and GSE70657 (fig. S1 and fig. S5H). scRNAseq_Dataset3/: Analysis pipeline for merging GSE59114 (C57BL/6), GSE59114 (DBA), and GSE70657 (fig. S1 and fig. S11F and G). Geneformer_codes/: Jupyter notebooks for fine-tuning tokenized data (6-layer and 12-layer models), in silico perturbation, and gene regulatory network inference using attention rollout; includes model safetensors and code for figure generation (Fig. 3). Python requirements.txt and setup.py for Geneformer v0.1.0 are also provided. The pretrained Geneformer model and template notebooks were obtained from the Hugging Face Hub (https://huggingface.co/ctheodoris/Geneformer), commit hash: c81f6f912c7e71483b9ca9560024904b3891f0bf Geneformer_network_visualization/: .graphml files for subgraphs containing HSC core genes, aging-associated HSC genes, and young HSC genes; includes an R script for network visualization and computation of graph features (Fig. 4H to K and Fig. 5I to M). FL_RNAseq_GSVA/: Gene set variation analysis files used to compute enrichment scores for Group 1 genes, Group 2 genes, and Young genes in fetal, neonatal, and young adult HSCs (fig. S1I). HSC_RNAseq/: Bulk RNA-seq data from freshly isolated and cultured young/old HSCs; includes scripts for alignment, quantification, differential expression analysis (DESeq2), fGSEA, and gene sets (Fig. 2 and fig. S2). HSC_ATACseq/: Bulk ATAC-seq data from freshly isolated and cultured young/old HSCs; includes scripts for alignment, peak counting, peak files, count matrices, differential accessibility analysis (edgeR), merged bigWig files, HOMER scripts, chromVAR scripts, and visualization scripts (Fig. 2 and fig. S2 and fig. S11B). Pbx1_RNAseq/: Bulk RNA-seq data from mock- or Pbx1-overexpressing young/old HSCs; includes scripts for alignment, quantification, differential expression analysis (DESeq2), fGSEA, and gene sets (Fig. 5 and fig. S6 and fig. S11A). Pbx1_ATACseq/: Bulk ATAC-seq data from mock- or Pbx1-overexpressing young/old HSCs; includes scripts for alignment, peak counting, peak files, count matrices, differential accessibility analysis (edgeR), merged bigWig files, HOMER scripts, TOBIAS footprinting scripts, chromVAR scripts, scripts for peak enrichment around promoters and enhancers for specified gene sets, and visualization scripts (Fig. 5 and fig. S7 and fig. S11B). TF_screening/: Data on CD48 and CD150 expression following transcription factor overexpression in HSCs, along with the visualization code, are provided (Fig. 4A). Pbx1KO_old_genes_overlap/: Computation of the statistical significance of the overlap between old HSC genes and Pbx1-regulated genes inferred from Pbx1 knockout data (Ficara F, et al. Cell Stem Cell. 2008; fig. S4F). Old_markers_after_culture_and_BMT/: Data and code for analyses of CD48 expression following bone marrow transplantation (fig. S5E and F) and expression levels of selected surface markers after 6 days of culture (fig. S5C). TF_OE_KD_qPCR/: Scripts and data are provided for qPCR analyses following transcription factor overexpression or knockdown in young HSCs, along with visualization (fig. S11D and E). Western_Blot_images/: Raw western blot images and visualization scripts (Fig. S5I). 5FU_PB_analysis/: Scripts and source data for peripheral blood analyses following treatment with 5-FU and a Pbx1 inhibitor (Fig. 6H). BMT_chimerism/: Scripts and source data for peripheral blood analyses following bone marrow transplantation (Fig. 6A to F and fig.S8 to 10).



