CRISPR-HAWK Data and Results
收藏资源简介:
Dataset Description This archive contains all input data and pre-computed outputs required to reproduce the analyses reported in the manuscript “CRISPR-HAWK: Haplotype- and Variant-Aware Guide Design Toolkit for CRISPR–Cas.” The dataset supports variant- and haplotype-resolved CRISPR guide RNA (gRNA) design and downstream functional interpretation. The directory is organized as follows: data/annotations/ Genomic feature annotations used to characterize candidate guide locations, including: COSMIC – catalogue of somatic cancer mutations DHS – DNase I hypersensitive sites ENCODE – experimentally annotated regulatory elements GENCODE – protein-coding gene annotations All files are provided in compressed and indexed BED format to enable efficient genomic interval queries. data/genomes/hg38/ Reference genome sequences for the human GRCh38/hg38 assembly, including: FASTA files (.fa) Index files (.fai) Only the chromosomes analyzed in the study are included. data/regions/ BED files defining the genomic loci supplied to CRISPR-HAWK for guide design: cas9.bed — regions targeted using SpCas9 (NGG PAM; 20-nt guides) cpf1.bed — regions targeted using Cas12a/Cpf1 (TTTV PAM; 23-nt guides) These loci include both therapeutic CRISPR targets and benchmark regions described in the manuscript. data/variants/ Population-scale human genetic variation datasets used to construct variant- and haplotype-resolved target sequences: 1000 Genomes Project (1000G) (VCFs and indexes downloaded using the helper script provided) HGDP (VCFs and indexes downloaded using the helper script provided) gnomAD v4.1 genomes(pre-processed using crisprhawk convert-gnomad-vcfs) CRISPR-HAWK integrates these datasets during guide discovery to account for genetic diversity across global populations. data/off-targets/ Off-target sites predicted using CRISPRme v2.1.7 (web interface) for: the therapeutic guide sg1617, and all alternative gRNAs identified by CRISPR-HAWK targeting the same locus. Search parameters: PAM: NGG Maximum mismatches: 6 Maximum bulges: 2 (DNA and RNA allowed) Variant datasets: 1000G + HGDP These files were used to evaluate how variant-matched guide design influences predicted off-target risk. results/ Pre-computed CRISPR-HAWK guide design outputs for all combinations of: CRISPR nucleases population variant datasets genomic target regions Outputs include variant- and haplotype-resolved guide candidates, annotations, and haplotype summaries. These tables can be used directly to regenerate the figures presented in the manuscript. results_nsamples/ Pre-computed CRISPR-HAWK outputs for all combinations of: * Cas nucleases * variant datasets * genomic regions with pre-computed number of samples. They can be used directly to regenerate the manuscript figures.



