Population-scale skeletal muscle single-nucleus multi-omic profiling reveals extensive context specific genetic regulation
收藏资源简介:
Data accompanying the manuscript "Population-scale skeletal muscle single-nucleus multi-omic profiling reveals extensive context specific genetic regulation". Filename: Description 1. snrna-cell-type-specific-genes.tsv: Normalized expression scores for genes in each cell-type cluster 2. eqtl_permutation_scan.tar.gz: # Permutation scan eQTL results in each cell-type cluster. For significant eQTLs, filter for qvalue < 0.05. Columns: # gene_name: gene name # start_pheno: phenotype start (gene TSS) # end_pheno: phenotype end (gene TSS) # strand: gene strand # n_variants_tested: number of variants tested for the gene # distance_var_pheno: distance of the variant with the gene TSS # snp: snp ID # snp_chrom: snp chromosomze # snp_start: snp start pos # snp_end: snp end pos # n_effective_tests: number of effective tests # p_nominal: nominal p value # slope: slope/beta of the linear regression. Keyed on the alt allele # se: standard error of the slope # p_beta: beta distribution adjusted p value # qvalue: qvalue (Storey) 3. caqtl_permutation_scan.tar.gz: # Permutation scan caQTL results in each cell-type cluster. For significant eQTLs, filter for qvalue < 0.05 Columns: # gene_name: peak feature coordinates # start_pheno: phenotype start (ATAC summit, i.e. peak feature midpoint) # end_pheno: phenotype end (ATAC summit, i.e. peak feature midpoint) # strand: strand ("." for ATAC) # n_variants_tested: number of variants tested for the peak feature # distance_var_pheno: distance between the variant with the peak feature # snp: snp ID # snp_chrom: snp chromosome # snp_start: snp start pos # snp_end: snp end pos # n_effective_tests: number of effective tests # p_nominal: nominal p value # slope: slope/beta of the linear regression. Keyed on the alt allele # se: standard error of the slope # p_beta: beta distribution adjusted p value # qvalue: qvalue (Storey) 4. eqtl_credible_sets.tar.gz: # eQTL credible set. The file name denotes the egene and the signal hit id. Bed file columns: # 1: snp chromosome # 2: snp start # 3: snp end # 4: snp chrom_pos_ref_alt # 5: Bayes Factor # 6: PIP # 7: SNP rsid 5. caqtl_credible_sets.tar.gz: # caqtl credible set. The file name denotes the egene and the signal hit id. Bed file columns: # 1: snp chromosome # 2: snp start # 3: snp end # 4: snp chrom_pos_ref_alt # 5: Bayes Factor # 6: PIP # 7: SNP rsid 6. cicero_all.tar.gz # Cicero coaccessibility results. Columns # Peak 1 : Macs2 narrowpeak coordinate for peak 1 # Peak 2 : Macs2 narrowpeak coordinate for peak 2 # coaccess: Cicero coaccessibility score 7. gene_peak_coaccessibility.tar.gz: Cicero coaccessibility results between peak and genes. Macs2 arrow peaks in the TSS+1kb upstream region are assigned that gene name. Columns # Peak 1 : Macs2 narrowpeak coordinate for peak 1 # gene_name: Assigned gene # Peak 2 : Macs2 narrowpeak coordinate for peak 2 # coaccess: Cicero coaccessibility score 8. coloc-eqtl-caqtl.tsv: # Summary of eQTL-caQTL coloc in each cluster. Columns: # nsnps: Number of SNPs in the region # eqtl_hit: SNP with the highest bayes factor in the SuSiE eQTL credible set # caqtl_hit: SNP with the highest bayes factor in the SuSiE caQTL credible set # PP.H0.abf: Coloc posterior probability for no signal # PP.H1.abf: Coloc posterior probability for signal in dataset 1 # PP.H2.abf: Coloc posterior probability for signal in dataset 2 # PP.H3.abf: Coloc posterior probability for different signal in datasets 1 and 2 # PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2 # idx1: Index of the SuSiE credible set for dataset 1 # idx2: Index of the SuSiE credible set for dataset 2 # cluster: cluster name # egene: eGene name # capeak: caPeak coordinates 9. cit-mrs-summary.tsv: Summary from CIT and MR Steiger directionality tests. Columns: # cluster: cluster name # egene: eGene name # capeak: caPeak coordinates # eqhit: SNP with the highest bayes factor in the SuSiE eQTL credible set # cahit: SNP with the highest bayes factor in the SuSiE caQTL credible set # p.cit_c_c-e: P value for CIT causal cahit-ca-to-e model # q.cit_c_c-e: q value for CIT causal cahit-ca-to-e model # p.cit_rc_c-e: P value for CIT reverse-causal eqhit-ca-to-e model # q.cit_rc_c-e: value for CIT reverse-causal eqhit-ca-to-e model # p.cit_c_e-c: P value for CIT causal eqhit-e-to-ca model # q.cit_c_e-c: q value for CIT causal eqhit-e-to-ca model # p.cit_rc_e-c: P value for CIT reverse-causal cahit-e-to-ca model # q.cit_rc_e-c: q value for CIT reverse-causal cahit-e-to-ca model # cit_direction: Direction inferred from CIT # correct_causal_direction--ca-to-e: MR Steiger directionality test - is ca-to-e direction correct? # correct_causal_direction--e-to-ca: MR Steiger directionality test - is e-to-ca direction correct? # sensitivity_ratio--ca-to-e: MR Steiger Sensitivity ratio for ca-to-e model # sensitivity_ratio--e-to-ca: MR Steiger Sensitivity ratio for e-to-ca model # steiger_test--ca-to-e: MR Steiger directionality test P value for ca-to-e model # steiger_test--e-to-ca: MR Steiger directionality test P value for e-to-ca model # steiger_q--ca-to-e: MR Steiger directionality test q value for ca-to-e model # steiger_q--e-to-ca: MR Steiger directionality test q value for e-to-ca model # mrs_direction: Direction inferred from MR Steiger # direction: Direction inferred requiring consistent results between CIT and MR Steiger directionality test 10. coloc-gwas-eqtl.tsv and 11. coloc-gwas-caqt.tav # Summary of e/caQTL coloc with GWAS in each cluster. Columns: # nsnps: Number of SNPs in the region # gwas_hit: SNP with the highest bayes factor in the SuSiE GWAS credible set # eqtl_hit: SNP with the highest bayes factor in the SuSiE eQTL credible set # caqtl_hit: SNP with the highest bayes factor in the SuSiE caQTL credible set # PP.H0.abf: Coloc posterior probability for no signal # PP.H1.abf: Coloc posterior probability for signal in dataset 1 # PP.H2.abf: Coloc posterior probability for signal in dataset 2 # PP.H3.abf: Coloc posterior probability for different signal in datasets 1 and 2 # PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2 # idx1: Index of the SuSiE credible set for dataset 1 # idx2: Index of the SuSiE credible set for dataset 2 # cluster: cluster name # egene: eGene name # capeak: caPeak coordinates # p12min: Min prior p12 where the PP H4 > 0.5. Lower this value, more robust is the colocalization # trait: GWAS trait name # gwas_locus: GWAS locus name for the coloc test - a 250kb left and right flanking genomic window on this SNP was considered for testing coloc between all pairs of GWAS/QTL signals identified in this region # traitname: Expanded GWAS trait name # variable_type: GWAS type # source: Source of GWAS - either UKBB or other study 12. supplementary_tables.xlsx: Supplementary tables from the manuscript. 13. eqtl_full_scan.tar.gz # Full eQTL cis scan, sorted by variant chrom:pos and tabix indexed. Columns: # snp_chrom: snp chromosome # snp_start: snp start # snp_end: snp end # snp: snp id # gene_name: gene name # chrom: chromosome # start_pheno: phenotype start (gene TSS) # end_pheno: phenotype end (gene TSS) # strand: gene strand # n_variants_tested: number of variant tested # distance_var_pheno: distance between the variant with the gene TSS # p_nominal: nominal p value # r2: the r squared of the linear regression # slope: slope/beta of the linear regression. Keyed on the alt allele # se: standard error of the slope # best_hit: whether this variant was the best hit for this phenotype.



