Pan-mammalian pan-cancer RNA-seq datasets processed with Paipu
收藏资源简介:
This dataset is associated with the manuscript "The Paipu framework enables creation of a large-scale mammalian cancer transcriptomics atlas". It contains pan-mammalian pan-cancer RNA-seq data processed using the Paipu pipeline. This data was compiled from NCBI SRA and processed consistently to generate gene expression data and harmonized sample metadata. There are two versions of the dataset: With duplicates: includes all samples after processing (Paipu_with_dups_*.tsv) Deduplicated: highly similar samples removed using sample expression correlation within species (Paipu_deduplicated_*.tsv) Each dataset includes sample metadata (*_metadata.tsv) and a gene expression matrix (*_expression.tsv). Gene expression values were normalized using trimmed mean of M-values (TMM) and transformed to log2 counts per million (CPM) with a prior count of 1.



