XChrom: a cross-cell chromatin accessibility prediction model integrating genomic sequences and cellular context
收藏资源简介:
XChrom is a deep learning framework that enables genome-wide cross-cell prediction of chromatin accessibility by integrating genomic sequence information with single-cell transcriptomics-defined cell identity within a unified model. Here, we provide processed data for the XChrom's tutorial and article. 1_within_sample.zip contains the trained model parameters for XChrom on the m_brain dataset. One can directly load E1000best_model.h5 as a trained model for subsequent analysis. 2_cross_samples.zip contains paired scRNA-seq and scATAC-seq data extracted from the BMMCs dataset for the s1d1 and s2d1 samples, stored as s1d1_s2d1.h5ad. We performed batch effect correction on both s1d1 scRNA-seq and s2d1 scRNA-seq using Harmony. The batch-corrected cell embeddings, typically named 'X_pca_harmony', were stored in the respective scRNA-seq datasets of both samples and used as input for model training and prediction. We saved them as train_rna_harmony.h5ad and test_rna_harmony.h5ad. Additionally, we provide the XChrom model parameters trained on the s1d1 dataset, where one can directly load E1000best_model.h5 as a trained model for further analysis. We used a publicly available single-cell multiome dataset collected from the motor cortex of human, macaque, marmoset, and mouse. Each species contains multiple samples. We extracted paired data and assigned the annotated cell types to prepare the dataset for our study. We integrated scRNA data from different samples within each species using Harmony and then performed cross-species integration based on 1-to-1 orthologous genes. The resulting batch-corrected cell embeddings are stored in RNA.obsm['X_pca_harmony']. The processed scRNA and scATAC paired data (after filtering and batch correction) have been uploaded to 3_cross_species.zip. We also provide the XChrom model parameters trained on the mouse (mop3c2) sample (mouse_trained model), where one can directly load E1000best_model.h5 as a trained model for further analysis. Additionally, we have uploaded the calculated TF activity matrices, which were inferred using the human_trained model on human samples (Hmodel_m1d1_tfactivity.h5ad) and mouse samples (Hmodel_M_tfactivity.h5ad), as well as the mouse_trained model on human samples (Mmodel_m1d1_tfactivity.h5ad) and mouse samples (Mmodel_M_tfactivity.h5ad). Furthermore, we have uploaded the ISM scores for human peak near PSEN1 that lacks homology to the mouse genome, calculated using the mouse_trained model (peak0_ism.npy), which can be used for visualization across different cell types and for investigating candidate TFs. 4_covid.zip contains the filtered PBMC10k training data (pbmc10k_atac.h5ad), as well as batch-corrected COVID-19 and PBMC10k scRNA-seq data (pbmc10k_rna_harmony.h5ad and covid_rna_harmony.h5ad). We also provide the XChrom model parameters trained on PBMC10k (E1000best_model.h5), along with the TF activity inferred for the COVID-19 data (tf_activity.h5ad). Source data can be downloaded with the links provided in the article or here. m_brain dataset: https://support.10xgenomics.com/single-cell-multiome-atac-gex/datasets/2.0.0/e18_mouse_brain_fresh_5k BMMCs dataset: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE194122 The motor cortex of human, macaque, marmoset, and mouse datasets: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE229169 COVID-19 dataset: https://nubes.helmholtz-berlin.de/s/wqg6tmX4fW7pci5/download



