遇见数据集

Population-scale skeletal muscle single-nucleus multi-omic profiling reveals extensive context specific genetic regulation

收藏
Zenodo2025-07-29 更新2026-05-26 收录
官方服务:

资源简介:

Data accompanying the manuscript "Population-scale skeletal muscle single-nucleus multi-omic profiling reveals extensive context specific genetic regulation". Note: For ATAC fragment files, e,caQTL full cis scan summary files, clustering objects, please see the CMDGA portal (https://cmdga.org/search/?searchTerm=stephen-parker%3AVarshney2024)For raw data including fastq files, please see dbGaP repo phs001048.v3.p1 Data in this repository includes: Filename: Description 1. list of 8,666 genes for which exon-only counts were considered. See methods section "Adjusting RNA counts for overlapping gene annotations" in the manuscript. 2. nucleus_sample_cluster_map.tsv: nucleus-sample-cluster map with other QC info. # index: nucleus identified syntax <modality>.<batch>.NM.<10X channel>.<barcode> # UMAP_1, UMAP_2: UMAP coordinates for visualization# modality: rna or atac# batch: processing batch identifier# hqaa_umi: high quality autosomal alignments (HQAA) for atac nuclei, unique molecular identifier (UMI) for tna # fraction_mitochondrial: fraction of reads mapping to the mitochondrial genome# cohort: sample cohort# tss_enrichment: TSS enrichment for atac nuclei# coarse_cluster_name: cluster name 3. peaks.tar.gz: snATAC peak features including:# consensus-summits.bed: consensus summits along with the cell type that the summits was highest in.# narrow peaks in clusters# consensus summit feature (summit +- 150bp) identified in each cluster - these were used in GWAS enrichments. 4. snrna-cell-type-specific-genes.tsv: Normalized expression scores for genes in each cell-type cluster 5. eqtl_permute.tar.gz: Permutation scan eQTL in each cell-type cluster. Columns: # variant: syntax <chrom>:<hg38 pos>:<ref>:<alt># effect_allele: effect allele (was the alt allele)# other_allele: non-effect allele# feature: gene name# featureCoordinates_tss: gene TSS# p-value: nominal p value# beta: slope/beta of the linear regression. Keyed on the alt allele# se: standard error of the slope# snp: SNP ID# strand: gene strand# n_variants_tested: number of variants tested for the gene# distance_var_pheno: distance of the variant with the gene TSS# n_effective_tests: number of effective tests# p_beta: beta distribution adjusted p value# qvalue: qvalue (Storey) 6. caqtl_permute.tar.gz: # Permutation scan caQTL in each cell-type cluster. Columns: # variant: syntax <chrom>:<hg38 pos>:<ref>:<alt># effect_allele: effect allele (was the alt allele)# other_allele: non-effect allele# feature: peak feature coordinates# p-value: nominal p value# beta: slope/beta of the linear regression. Keyed on the alt allele# se: standard error of the slope# snp: SNP ID# n_variants_tested: number of variants tested for the gene# distance_var_pheno: distance of the variant with the gene TSS# n_effective_tests: number of effective tests# p_beta: beta distribution adjusted p value# qvalue: qvalue (Storey) 7. eqtl_credible_sets.tar.gz: # eQTL credible set. The file name denotes the egene and the signal hit id. Bed file columns: # 1: snp chromosome# 2: snp start# 3: snp end# 4: snp chrom_pos_ref_alt# 5: Bayes Factor # 6: PIP# 7: SNP rsid 8. caqtl_credible_sets.tar.gz: # caqtl credible set. The file name denotes the capeak and the signal hit id. Bed file columns: # 1: snp chromosome# 2: snp start# 3: snp end# 4: snp chrom_pos_ref_alt# 5: Bayes Factor # 6: PIP# 7: SNP rsid 9. cicero_all.tar.gz # Cicero coaccessibility results. Columns# Peak 1: Macs2 narrowpeak coordinate for peak 1# Peak 2: Macs2 narrowpeak coordinate for peak 2# coaccess: Cicero coaccessibility score 10. cicero_gene_tss.tar.gz: Cicero coaccessibility results between peak and genes. Macs2 narrow peaks in the TSS+1kb upstream region are assigned that gene name. Columns# Cicero coaccessibility results between peak and genes. Macs2 narrow peaks in the TSS+1kb upstream region are assigned that gene name.Columns# Peak 1: Macs2 narrowpeak coordinate for peak 1# gene_name: Assigned gene# Peak 2: Macs2 narrowpeak coordinate for peak 2# coaccess: Cicero coaccessibility score## Peak1 is the narrowpeak in the TSS region, peak2 is the distal peak 11. mash.tar.gz Mashr results for e/caQTL - lfsr, posterior means and posterior SD for each tested eSNP-eGene, caSNP-caPeak pair. 12. cellregmap.tar.gz: Cellregmap results for endothelial nucleus-level eQTL scans.## Persistent genetic effect beta_g was calculated in a simple association model. ## An interaction model was fit to test for GxC effect. columns:# rho1, g2, e1, and eps2 are variance component measures outputs from CellRegMap corresponding to interaction, genetic, environment and residual variance components. # p_nominal: nominal p from cellRegMap# kind: model kind in CellRegMap - simple association or interaction# beta_g: Persistent genetic effect# gene_name: gene name for eQTL or peak feature name for caQTL# context: context used either factors (continuous) or subclusters (discrete)# snp: index snp for which model is fit. This is the most significant identified snp from our standard e,caQTL scans. chrom-hg38pos-rsid 13. coloc-eqtl-caqtl.tsv: # Summary of eQTL-caQTL coloc in each cluster. Columns:# nsnps: Number of SNPs in the region# eqtl_hit: SNP with the highest Bayes factor in the SuSiE eQTL credible set# caqtl_hit: SNP with the highest Bayes factor in the SuSiE caQTL credible set# PP.H0.abf: Coloc posterior probability for no signal# PP.H1.abf: Coloc posterior probability for signal in dataset 1# PP.H2.abf: Coloc posterior probability for signal in dataset 2# PP.H3.abf: Coloc posterior probability for different signals in datasets 1 and 2# PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2# idx1: Index of the SuSiE credible set for dataset 1# idx2: Index of the SuSiE credible set for dataset 2# cluster: cluster name# egene: eGene name# capeak: caPeak coordinates 14. cit-mrs-summary.tsv: Summary from CIT and MR Steiger directionality tests. Columns:# cluster: cluster name# egene: eGene name# capeak: caPeak coordinates# eqhit: SNP with the highest Bayes factor in the SuSiE eQTL credible set# cahit: SNP with the highest Bayes factor in the SuSiE caQTL credible set# p.cit_c_c-e: P value for CIT causal cahit-ca-to-e model# q.cit_c_c-e: q value for CIT causal cahit-ca-to-e model# p.cit_rc_c-e: P value for CIT reverse-causal eqhit-ca-to-e model # q.cit_rc_c-e: value for CIT reverse-causal eqhit-ca-to-e model # p.cit_c_e-c: P value for CIT causal eqhit-e-to-ca model# q.cit_c_e-c: q value for CIT causal eqhit-e-to-ca model# p.cit_rc_e-c: P value for CIT reverse-causal cahit-e-to-ca model # q.cit_rc_e-c: q value for CIT reverse-causal cahit-e-to-ca model # cit_direction: Direction inferred from CIT # correct_causal_direction--ca-to-e: MR Steiger directionality test - is ca-to-e direction correct?# correct_causal_direction--e-to-ca: MR Steiger directionality test - is e-to-ca direction correct?# sensitivity_ratio--ca-to-e: MR Steiger Sensitivity ratio for ca-to-e model # sensitivity_ratio--e-to-ca: MR Steiger Sensitivity ratio for e-to-ca model# steiger_test--ca-to-e: MR Steiger directionality test P value for ca-to-e model# steiger_test--e-to-ca: MR Steiger directionality test P value for e-to-ca model# steiger_q--ca-to-e: MR Steiger directionality test q value for ca-to-e model# steiger_q--e-to-ca: MR Steiger directionality test q value for e-to-ca model# mrs_direction: Direction inferred from MR Steiger# direction: Direction inferred requiring consistent results between CIT and MR Steiger directionality test 15. coloc-gwas-eqtl.tsv and16. coloc-gwas-caqtl.tsv # Summary of e/caQTL coloc with GWAS in each cluster. Columns:# nsnps: Number of SNPs in the region# gwas_hit: SNP with the highest bayes factor in the SuSiE GWAS credible set# eqtl_hit: SNP with the highest bayes factor in the SuSiE eQTL credible set# caqtl_hit: SNP with the highest bayes factor in the SuSiE caQTL credible set# PP.H0.abf: Coloc posterior probability for no signal# PP.H1.abf: Coloc posterior probability for signal in dataset 1# PP.H2.abf: Coloc posterior probability for signal in dataset 2# PP.H3.abf: Coloc posterior probability for different signal in datasets 1 and 2# PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2# idx1: Index of the SuSiE credible set for dataset 1# idx2: Index of the SuSiE credible set for dataset 2# cluster: cluster name# egene: eGene name# capeak: caPeak coordinates# p12min: Min prior p12 where the PP H4 > 0.5. Lower this value, more robust is the colocalization# trait: GWAS trait name# gwas_locus: GWAS locus name for the coloc test - a 250kb left and right flanking genomic window on this SNP was considered for testing coloc between all pairs of GWAS/QTL signals identified in this region # traitname: Expanded GWAS trait name# variable_type: GWAS type # source: Source of GWAS - either UKBB or other study 17. supplementary_tables.xlsx: Supplementary tables from the manuscript.Information included in sheets:1. "marker_genes": Marker genes known from literature used to annotate clusters2. "n_nuclei": n pass-QC nuclei per modality-sample-cluster 2. "snrna_GO_enrichment": GO term enrichment: matrix of cluster vs top 2 GO terms 3. "qtl_scan_info": e/caQTL scan infocluster: clusterntested_eqtl: N genes tested for eQTLnsig_eqtl: N significant (5% FDR) eGenesn_pheno_pcs_eqtl: N phenotype PCs considered for eQTLratio_eqtl: Ratio of N eGenes/N genes testednsig_caqtl: N peaks tested for caQTLntested_caqtl: N significant (5% FDR) caPeaksn_pheno_pcs_caqtl: N phenotype PCs considered for caQTLratio_caqtl: Ratio of N caPeaks/N peaks testednsamples_eqtl: N samples for eQTLnsamples_caqtl: N samples for caQTL 4. "gwas_trait_list": GWAS trait infotrait: GWAS trait IDtraitname: GWAS trait descriptionvariable_type: GWAS type. case/control (cc), continuous_irnt=continuous inverse-normal transformedsource: GWAS sourcedoi: GWAS study DOI 5. "traits_in_ldsc_baseline" - list of annotations included in the baseline model for LDSC 6. "gwas_enrichment_in_peaks" GWAS enrichment in cluster peaks (S-LDSC) 7. "gwas_enrichment_in_qtl_peaks" GWAS enrichment in QTL peaks (fGWAS) # fGWAS results comparing GWAS enrichment in type 1 annotationsCI_lower_ln, estimate_ln, CI_upper_ln: natural log of lower confidence interval, estimate, and upper confidence intervaltrait: trait idtraitname: trait nameannotation: annotationsig: 1 if CIs don't overlap 0, otherwise 0 8. t2d_gwas_caqtl_coloc and9. t2d_gwas_eqtl_coloc:Summary of e,caQTL coloc with T2D GWAS in each cluster, along with target gene nominations. Columns:nsnps: Number of SNPs in the regiongwas_hit: SNP with the highest bayes factor in the SuSiE GWAS credible seteqtl_hit: SNP with the highest bayes factor in the SuSiE eQTL credible setcaqtl_hit: SNP with the highest bayes factor in the SuSiE caQTL credible setPP.H0.abf: Coloc posterior probability for no signalPP.H1.abf: Coloc posterior probability for signal in dataset 1PP.H2.abf: Coloc posterior probability for signal in dataset 2PP.H3.abf: Coloc posterior probability for different signal in datasets 1 and 2PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2idx1: Index of the SuSiE credible set for dataset 1idx2: Index of the SuSiE credible set for dataset 2cluster: cluster nameegene: eGene namecapeak: caPeak coordinatesp12min: Min prior p12 where the PP H4 > 0.5. Lower this value, more robust is the colocalizationtrait: GWAS trait iddiamante_gwas_locus: GWAS signal from the DIAMANTE 2018 study. Some signals that our SuSiE runs identified were not present in the original study in which case this column is NAtraitname: Expanded GWAS trait namecapeak_in_tss: caPeak in TSS + 1kb upstream region of a genegene_target_standard_cicero: caPeak coaccessible with TSS peak of a gene considering nuclei from all samples for co-accessibilitygene_target_allelic_cicero: caPeak coaccessible with TSS peak of a gene considering nuclei from samples homozygous for the caSNP allele associated with increased accessibilitygwashit_nominal_egene: gwas_hit nominally associated with these genes nominated in the columns capeak_in_tss, gene_target_standard_cicero, and gene_target_allelic_cicero 10. MPRA results for the C2CD4A locus

提供机构:
Zenodo
创建时间:
2025-07-29
二维码
社区交流群
二维码
科研交流群
商业服务