Simulated data used in the KuPID publication
收藏资源简介:
This dataset contains the simulated transcript annotations and sequencing data used in 'K-mer-based Upstream Preprocessing of long reads for Isoform Discovery'. All simulated reads were based on HiFi sequencing of GENCODE's v48 human transcriptome. reduced_annotations: Set of novel annotations created via the 'reduction' method. A given percentage of isoforms (20%-80%) were randomly selected to be treated as novel. In addition, it contains an altered version of GENCODE's v48 human transcriptome where the novel isoforms have been omitted. Contains 4 replicates. yasim_annotations: Set of novel annotations created via the YASIM software (Su et al. 2024). Contains 4 replicates. reduced_data: Simulated HiFi reads based on the reduction method annotations. Reads are provided as a bam file, and were mapped to release 114 of the Ensembl human genome with the minimap2 aligner. Total number of reads varied from 6 million (20% of the isoforms are novel) to 1.5 million (80% of the isoforms are novel). Contains 4 replicates. yasim_data: Simulated HiFi reads based on the YASIM method annotations. Reads are provided as a bam file, and were mapped to release 114 of the Ensembl human genome with the minimap2 aligner. Total number of reads varied from 4 million (20% of the isoforms are novel) to 1 emillion (80% of the isoforms are novel). Contains 4 replicates.



