Simulation data for benchmarking de novo long read transcriptome assembly software
收藏资源简介:
Method of simulation of differentially expressed biological replicates We first obtained a subset of transcripts that are widely expressed in the GTEx v9 dataset using Gencode.v44 annotation (18145 genes, 40509 transcripts), and stored their count per million (CPM) values as CTRL baseline. We manually changed transcript expression in 3 scenarios: (1) select 1000 genes and change all transcripts belonging to that gene concordantly (500 genes 2 fold up and 500 genes 2 fold down), (2) select 1000 genes, and then select 2 random transcripts from the gene and swap their expression, (3) select 1000 genes, and then select 1 random transcript to change its expression (500 transcripts 2 fold up and 500 transcripts 2 fold down). The updated counts and CPM were stored as DE baseline. We then generated count matrix and count per million (CPM) matrix for 3 CTRL replicates and 3 DE replicates with gamma distribution, followed by a Poisson distribution (Baldoni et al., 2023). Both long-read and short-read FASTQ files were simulated using SQANTI-SIM (v 0.2.1) (Mestre-Tomás et al., 2023). The long read data contains 6 million reads in total, and an average read length of 1085 bp, and short read data is 100 bp paired-end. We then subsampled the short-read data to match the total base pair used in the Nanopore data (6.5 billion bases). This simulated data is non-stranded, and contains 2000 DGEs, 2000 DTU-genes, 5927 DTU-transcripts and 6933 DTEs.



