"A unified genetic perturbation language for human cellular programming" accompanying data
收藏资源简介:
Description of data: crispr_all_seq_seurat_object.rds: CRISPR-All-seq Seurat object (package version 5.2.0) containing all cells assigned to a single perturbation. Metadata includes replicate (10x lane), donor (n=3), construct (assigned using deMULTIplex2; see paper methods for details), and class (type of assigned perturbation; Cholla=KO, PricklyPear=synthetic gene, Saguaro=gene overexpression, Torch=shRNA gene KD, or Kingcup=alternative CAR intracellular signaling domain). pacbio_basescreen_bc_and_construct_assignments.csv: Three-column CSV containing identified barcode and construct for each read. Column names are ID (read), barcode, and construct. Briefly, barcodes are detected via a brute-force with 0 tolerated hamming distance and constructs are detected based on alignment to a custom reference using minimap2. pacbio_comboscreen_bc_and_construct_assignments.csv: Ten-column CSV containing identified barcodes and constructs for each read. The columns titled signaling_bc, gene_bc, triplex_bc, ko_bc, and kd_bc contain the detected barcode for the respective perturbation type for that read (Key column contains the read ID). There is only one possible triplex barcode, called triplex, but it's presence indicates proper cloning/inclusion of the triplex sequence in the 3' UTR which stabilizes the transcript following gRNA and shRNA processing which separates the coding sequence from the stabilizing 3' polyA tail. The ko, kd, gene, and signaling columns indicate which construct is detected by alignment for each respective perturbation type. In short, 4 separate custom references are constructed for each perturbation type, and minimap2 is used to align reads to those references (see paper methods for more detail). salmon.merged.gene_counts.csv: Bulk RNA-seq counts matrix generated via pseudoalignment with salmon, implemented within the nf-core/rnaseq pipeline. Rows contain genes and columns contain unique combinatorial constructs. base_screen_amplicon_counts_matrix.csv: Counts matrix generated from amplicon sequencing of the base CACTUS library following a repetitive stimulation assay. Each column represents a sample (donor & timepoint) and each row represents a construct in the library. combo_screen_amplicon_counts_matrix.csv: Counts matrix generated from amplicon sequencing of the combinatorial CACTUS library following a repetitive stimulation assay. Each column represents a sample (donor & timepoint) and each row represents a construct in the library. The construct names are in the following format: "KD_KO_Gene_Signaling."



