Raw and processed ONT long read sequencing data for "Complete genome sequences of four Ochrobactrum pituitosum strains and three Pseudochrobactrum saccharolyticum strains isolated from Caenorhabditis elegans gut microbiomes"
收藏资源简介:
Uploaded data contain raw and processed data from ONT sequencing of 8 samples from the gut microbiome of C. elegans: 5 isolates of Ochrobactrum pituitosa and 3 isolates of Pseudochrobactrum saccharolyticum. Details about the data and results are given in the Genome Announcement paper titled "Complete genome sequences of four Ochrobactrum pituitosum strains and three Pseudochrobactrum saccharolyticum strains isolated from Caenorhabditis elegans gut microbiomes". Bacterial Genome Assembly - Analysis Method Raw nanopore sequencing reads are assessed for quality and filtering. Filtlong v0.2.1 [1] is used to remove short and low-quality reads from the raw nanopore sequencing data. The high-quality nanopore sequencing reads are then used for de novo assembly of the bacterial genome. The assembly is performed using Flye v2.9.3 [2] with parameters optimized for bacterial genomes. The resulting contigs are further polished using Medaka v1.8 [3] to improve base accuracy. The assembled genomes are annotated using Bakta v.1.8.2 [4]. The annotation includes the prediction of coding sequences, tRNAs, rRNAs, and other genomic features based on published databases such as RefSeq and UniProt. The quality of the assembled genomes is assessed using various quality assessment tools (QUAST v5.2 [5], CheckM2 v1.0.1 [6], Mash v2.3 [7]). Genome completeness, contiguity, and accuracy are evaluated to ensure the reliability of the assemblies. A final QC step is included to ensure the purity of samples, which uses minimap2 v2.24 [8] to map the sequence-cleaned reads onto the assembly and finally employs Clair3 v1.0.4 [9] to call variations (SNPs and INDELS) within the assembled genome. Any variants detected along with the location in the bacterial assembly are reported. References 1. Filtlong: Ryan R Wick, https://github.com/rrwick/Filtlong2. Flye: Kolmogorov, Mikhail, Jeffrey Yuan, Yu Lin, und Pavel A. Pevzner. "Assembly of Long, Error-Prone Reads Using Repeat Graphs". Nature Biotechnology 37, Nr. 5 (Mai 2019): 540–46. https://doi.org/10.1038/s41587-019-0072-8.3. Medaka: ONT Research, https://github.com/nanoporetech/medaka4. Bakta: Schwengers, Oliver, Lukas Jelonek, Marius Alfred Dieckmann, Sebastian Beyvers, Jochen Blom, und Alexander Goesmann. "Bakta: rapid and standardized annotation of bacterial genomes via alignment-free sequence identification". Microbial Genomics 7, Nr. 11 (5. November 2021): 000685. https://doi.org/10.1099/mgen.0.000685.5. Quast: Gurevich, Alexey, Vladislav Saveliev, Nikolay Vyahhi, und Glenn Tesler. "QUAST: Quality Assessment Tool for Genome Assemblies". Bioinformatics (Oxford, England) 29, Nr. 8 (15. April 2013): 1072–75. https://doi.org/10.1093/bioinformatics/btt086.6. CheckM2: Chklovski, Alex, Donovan H. Parks, Ben J. Woodcroft, und Gene W. Tyson. "CheckM2: A Rapid, Scalable and Accurate Tool for Assessing Microbial Genome Quality Using Machine Learning". Nature Methods 20, Nr. 8 (August 2023): 1203–12. https://doi.org/10.1038/s41592-023-01940-w.7. Mash: Ondov, Brian D., Todd J. Treangen, Páll Melsted, Adam B. Mallonee, Nicholas H. Bergman, Sergey Koren, und Adam M. Phillippy. "Mash: Fast Genome and Metagenome Distance Estimation Using MinHash". Genome Biology 17, Nr. 1 (20. Juni 2016): 132. https://doi.org/10.1186/s13059-016-0997-x.8. Minimap2: Li, Heng. "Minimap2: Pairwise Alignment for Nucleotide Sequences". Bioinformatics (Oxford, England) 34, Nr. 18 (15. September 2018): 3094–3100. https://doi.org/10.1093/bioinformatics/bty191.9. Clair3: Zheng, Zhenxian, Shumin Li, Junhao Su, Amy Wing-Sze Leung, Tak-Wah Lam, und Ruibang Luo. "Symphonizing Pileup and Full-Alignment for Deep Learning-Based Long-Read Variant Calling". Nature Computational Science 2, Nr. 12 (Dezember 2022): 797–803. https://doi.org/10.1038/s43588-022-00387-x.



