Supplementary data for the publication in Clinical Epigenetics, <b>Allelic Expression Patterns of Imprinted and Non-imprinted Genes in Cancer Cell Lines from Multiple Histologies. </b>Authors: Julia Krushkal, Travis L. Jensen, George Wright, Yingdong Zhao. DOI: 10.1186/s13148-025-01883-3
收藏资源简介:
Supplementary dataset for the manuscript <b>Allelic expression patterns of imprinted and non-imprinted genes in cancer cell lines from multiple histologies</b><b> </b>by Julia Krushkal, Travis L. Jensen, George Wright, and Yingdong Zhao.Allelic expression patterns are summarized for 108 cancer cell lines from 9 pediatric and adult tumor histologies. Files are in compressed gzip format. The dataset includes 87 total data files. They include 81 files with a file for each analysis level (gene, isoform, and exon), cancer category (acute myeloid leukemia, bladder, breast, colorectal, head and neck, neuroblastoma, ovarian, pancreatic, and small cell lung cancer), and data type (count [Feature annotations; number of heterozygous SNVs, number of HTSeq read counts, and log10 feature length normalized number of HTSeq read counts for both RNA-seq and WES data; copy number]; summary [Feature annotations; feature algorithm assignment counts across cell lines]; and cutoff [Feature annotations; per-cell-line and feature algorithm assignment]). Additionally, 6 summary and cutoff files are provided for the pancancer dataset including all 108 cell lines at the gene, isoform, and exon levels.The source code for the pipeline for generating the data is provided in the archive <b>Source_code.zip</b>. The file <b>Genes assigned to multiple chromosomes.xlsx</b> provides the list of genes with reported assignments to multiple chromosomes, since their ambiguous mapping may introduce errors in their copy number and allelic expression inference.Gene and exon annotations are provided according to GENCODE, using lifted annotations from V38lift37 (Ensembl 104) mapped to hg19 (gencode.v38lift37).<b>WES:</b><b> </b>whole exome sequencing<b>Exon_ID, Transcript_ID, Gene_ID</b><b> </b>provide ID of a feature (exon, isoform, and gene)The “<b>_1</b>” , which is a part of each feature ID, is included in the GENCODE GTF/GFF feature annotations for all genes/transcripts/exons, corresponding to the mapping version “lift” from GRCh38 to GRCh37/hg19.<b>Gene_name,</b><b> </b>provides gene name<b>Gene_label</b>: <b>0</b><b> </b>- imprinted, <b>11</b><b> </b>- positive control, <b>12</b><b> </b>- negative control, <b>2</b><b> </b>- other<b>Chr:</b><b> </b>chromosome<b>Start, End:</b> genomic coordinates of the start and end of feature<b>Exon_IDs</b> and <b>Exon_ranges</b>: identifiers and boundaries of all exons which are annotated to be a part of a given gene<b>Overlapping_exon_bases</b>: the combined length of all isoforms in a given gene, some of which may be overlapping<b>Unique_exon_bases</b>: the total length of unique isoform bases in a given gene, with duplicate positions excludedIn the count data files, the following information is provided for each cell line and each feature:<b>COPY_NUMBER:</b><b> </b>gene level copy number<b>RNA_VCF_HET:</b><b> </b>number of heterozygous SNVs in the RNA-seq data<b>WES_VCF_HET:</b><b> </b>number of heterozygous SNVs in WES data<b>RNA_HTSEQ_COUNT</b>S: number of HTSeq read counts in RNA-seq data (measure of feature expression)<b>LOG10_LENGTH_NORM_RNA_HTSEQ_COUNTS</b>: log10(number of HTSeq read counts in RNA-seq data normalized by feature length)<b>WES_HTSEQ_COUNT</b>S: number of HTSeq read counts in WES data<b>LOG10_LENGTH_NORM_WES_HTSEQ_COUNTS</b>: log10(number of HTSeq read counts in WES data normalized by feature length)Each cutoff file provides inferred allelic expression patterns of a given feature in individual cell lines.Each summary file provides summary counts of a given feature in individual cell lines:<b>Cell_lines_with_copy_number_loss, Cell_lines_with_no_SNV_data, Cell_lines_with_no_expression. Cell_lines_with_biallelic_expression, Cell_lines_with_monoallelic_expression,</b><b> </b>and <b>Cell_lines_with_unresolved_status</b><b> </b>provide the number of cell lines with a particular expression pattern for a given feature.<b>Discrepancy_RNA_VCF_vs_WES_VCF</b> indicates cases where only RNA-seq but not WES data that indicate the presence of one or more heterozygous SNVs.



