Four benchmark datasets: Hsapiens 18S rRNA, curlcakes, artificial oligos and artificial oligos mixtures
收藏资源简介:
Benchmark datasets accompanying the article "Nanopore based computational method for detecting a wide range of epigenetic/epitranscriptomic modifications". 1. H. sapiens 18S rRNA dataset was generated from the HCT116 human colon carcinoma cell line. Dataset contains six directories corresponding to coverage depths of 10, 50, 75, 100, 200, and 500. Each directory includes Ivt directory (in vitro transcripts, no modified sites) and Native directory (native reads with modifications present). The number of reads in each subdirectory corresponds to its parent directory label (ranging from 10 to 500). All reads were selected from the original dataset (Naarmann-de Vries et al.) to approximately span the full reference sequence, resulting in approximatelly-uniform coverage depth matching the directory name (e.g. 10, 50,...500; see example) For details about dataset construction either see Vujaklija et al., or drop me an email at ivan.vujaklija@gmail.com. The original sequencing was performed by Naarmann-de Vries et al. using the R9.4.1 pore. Native reads were basecalled with Guppy v5.1.13 and ivt reads with Guppy v6.4.6. (We retained the original basecalls to reflect the publicly released data). All reads were resquiggled with Tombo v1.5. 2. Curlcake_dataset contains nanopore reads from four artificial molecules–Curlcake 1, Curlcake 2, Curlcake 3 and Curlcake 4 (2329, 2543, 2678, and 2795 bp in length respectively). These molecules were synthesized by Liu et al. using CURLCAKE software and subsequently transcribed in vitro with m6ATP, resulting in incorporation of m6A at all adenosine positions. The dataset comprises eight directories corresponding to coverage depths of 10, 50, 75, 100, 200, 500, 1000 and 2000. Each directory contains four CCx_test and four CCx_control(x=1-4) subdirectories. The number of reads in each subdirectory corresponds to its parent directory label. All reads were selected from the original dataset to approximately span the full reference sequences, resulting in approximatelly uniform coverage depth matching the directory name (e.g. 10, 50, 75,..,2000; see example). For details about dataset construction either see Vujaklija et al., or drop me an email at ivan.vujaklija@gmail.com. The original sequencing was performed by Liu et al. using the R9.4.1 pore. Reads were basecalled with Albacore v2.1.7, and resquiggled with Tombo v1.5. 3. Oligo_dataset contains nanopore reads of the same (~100nt in length) nucleotide sequence in four variants: oligo-1 (three m6A sites), oligo-2 (one I, one m5C, and one Ψ site), oligo-3 (one m62A, one m1G and one Am site) and oligo-control (modification free). This dataset was constructed from the original dataset published by Leger et al. The dataset comprises three independent samples: Sample-01, Sample-02 and Sample-03. Each sample directory contains eight subdirectoreis corresponding to different coverege depths of 10, 50, 75, 100, 200, 500, 1000, and 2000. Each of these, contain four subdirectories: oligo-var1, oligo-var2, oligo-var3 and oligo-control. For example, Sample-1/coverage-depth-0010/oligo-var1 subdirectory contains 10 reads, each containing 3 m6A modified sites. All reads were selected from the original dataset to approximately span the full reference sequences, resulting in approximatelly uniform coverage depth matching the directory name (e.g. 10, 50, 75,..,2000; see example). For details about dataset construction either see Vujaklija et al., or drop me an email at ivan.vujaklija@gmail.com. The original sequencing was performed by Leger et al. using the R9.4.1 pore, reads were basecalled with Guppy v3.2.10, and resquiggled with Tombo v1.5. 4. Oligo-mixtures_dataset contains the same four oligo variants as the Oligo_dataset. The dataset was constructed from the Oligo_dataset by in silico replacement of a fraction of modified reads with unmodified reads. Its purpose is to test algorithms in settings where only a fraction of the test reads (25%, 50%, or 75%) are modified at a particular site. Note that control samples contain only unmodified reads. The dataset comprises 10 independent samples, each with eight coverage depth directories (20, 50, 75, 100, 200, 500, 1000, 2000). Within each coverage depth directory, there are three stoichiometry subdirectories (25%, 50%, 75%), each containing four oligo directories (oligo-var1, oligo-var2, oligo-var3, oligo-control). In these, oligo-var1/2/3 directories contain the specified fraction of modified reads, while oligo-control directory contains only unmodified reads (e.g. Sample-1/Coverage depth 100/Stoichiometry 50% includes oligo-var1/2/3 directories with 50 modified and 50 unmodified reads, and oligo-control directory with 100 unmodified reads). For details about dataset construction either see Vujaklija et al., or drop me an email at ivan.vujaklija@gmail.com. The original sequencing was performed by Leger et al. using the R9.4.1 pore, reads were basecalled with Guppy v3.2.10, and resquiggled with Tombo v1.5. If you have any questions or would like additional clarification, please feel free to contact me at ivan.vujaklija@gmail.com, I’ll be happy to help.



