遇见数据集

Synthetic dataset of FASTQ files used to benchmark the rnaends R package performances

收藏
Zenodo2025-11-14 更新2026-05-26 收录
官方服务:

资源简介:

Synthetic dataset of FASTQ files used to benchmark the rnaends R package performances. The rnaends package is available at https://gitlab.com/rnaends/rnaends FASTQ files with variable sequencing depths (1 to 20 millions of reads per experimental condition) were generated based on observed distributions of 5’ ends in two conditions (5’PPP only RNAs vs. background noise as in the TSS-EMOTE protocol) to evaluate the performances of different features of the rnaends package: pre-processing of raw reads only validate then extract the mappable part of reads. checking the presence and validity of some sequences, i.e. a recognition sequence (CGGCACCAACCGAGG) and a control sequence (CGC) at precise locations on the reads (see the read structure below for details). checking the validity of UMIs and extracting them for further use for PCR duplicates removal. using barcodes for demultiplexing experimental conditions. quantification mapping count table downstream TSS 100 nucleotide long reads were generated. UMIs are 17 nucleotides long and are composed of only A, C and G. 12 barcodes were used corresponding to two experimental conditions with 3 or 6 replicates depending on the fact that the first 2 or 3 nucleotides of the barcodes are used: AAxx, ACxx, ATxx for 3 replicates of the 5’PPP RNAs condition, and AAAx, AACx, ACAx, ACCx, ATAx, ATCx for the 6 replicates CAxx, CCxx, CTxx for the 3 replicates of the background noise condition, and CAAx, CACx, CCAx, CCCx, CTAx, CTCx for the 6 replicates. As a result, benchmarks for 12 barcodes process 12 times 1-5-10-15-20 million reads per barcode/condition, and benchmarks with 6 barcodes process 2-10-20-30-40 million reads per barcode/condition. Thus, the biggest analysis presented here (6 barcodes with 20 millions reads per condition), consists in 240 million reads.

提供机构:
Zenodo
创建时间:
2025-11-14
二维码
社区交流群
二维码
科研交流群
商业服务