遇见数据集

Table S1. Sequencing summary, taxonomic classification, and coverage metrics for Schistosoma mansoni miracidia samples using Nanopore adaptive sampling.

收藏
Zenodo2025-04-28 更新2026-05-26 收录
官方服务:

资源简介:

Title:How useful is Nanopore adaptive sampling for sequencing Schistosoma mansoni miracidia? Overview:This dataset contains results from post-sequencing analysis used to evaluate the efficiency of adaptive sampling for enriching Schistosoma mansoni DNA and the taxonomic classification of DNA fragments using various bioinformatics tools. File(s): Results_sup.csv: Main data file containing numeric and categorical results for each sample processed. The file is comma-separated and includes both pre- and post-sequencing data, as well as taxonomic and quality control metrics. Methodology Summary:Sequencing reads were basecalled using Dorado (v0.8.1) with a super accurate model. Adaptive sampling was used to enrich for S. mansoni DNA. Reads marked with "stop_receiving" were considered enriched. Matching read IDs were searched for in Dorado basecalls to determine enrichment rates. Statistical tests (Shapiro-Wilk, t-test, Mann-Whitney U) were used to assess differences between washed and unwashed sample groups. For taxonomic identification, Kraken2 was run against a custom database including all RefSeq bacterial, viral, protozoal genomes, the human genome, the UNI_vec_core, and the V10 S. mansoni reference. Sequences identified as S. mansoniby Kraken2 but not adaptive sampling were BLASTed against a custom NCBI Schistosoma database (taxonomy code 6181) using BLAST+ (v2.9.0). Positive hits were realigned to the S. mansoni reference using MiniMap2 (v2.28). Quality-filtered reads (length >150 bp, mean Q >12) were mapped using MiniMap2, and Samtools (v1.21) was used to assess coverage depth and breadth. Vocabulary / Column Description: Column Group Description Sample number Numerical sample identifier Status Whether the sample was “Washed” or “Unwashed” prior to sequencing Pre sequencing Statistics before sequencing, e.g., mean read length Post sequencing Number and mean length of reads post-sequencing Adaptive sampling results Includes % enriched, number of reads enriched, and mean length of enriched reads Kraken2 results Percentage of S. mansoni DNA identified taxonomically by Kraken2 MiniMap results Number of reads aligned and coverage statistics after filtering (length >150 bp, Q >12) Samtools coverage and depth Number of mapped bases, genome breadth, and mean depth of coverage BLAST results Number of sequences with hits to various Schistosoma species including S. haematobium, S. rodhaini, S. spindale, and S. mansoni Some columns may contain ranges (e.g., "211:166896") referring to min:max basepair lengths. Usage Instructions:To replicate or reuse the analysis: Use Dorado (v0.8.1) for basecalling using the super accurate model. Determine enrichment based on “stop_receiving” tags in the adaptive sampling report. Use Kraken2 (with 0.1 confidence threshold) on a custom database as described. Extract sequences not enriched by adaptive sampling but classified as S. mansoni by Kraken2 using seqtk. Perform BLAST (v2.9.0+) against a custom NCBI Schistosoma database. Use MiniMap2 (v2.28) to map filtered reads (Nanofilt v2.8.0, Q>12, bp>150) - both enriched and those identified by Kraken2. Assess mapping statistics and coverage using Samtools (v1.21). Statistical analysis was performed in R (v4.4.2). Notes: The CSV includes summary-level data; The raw Nanopore sequencing data (FASTQ files) have been deposited in the NCBI Sequence Read Archive (SRA) and are associated with the following BioSample accessions: SAMN47928699, SAMN47928700, SAMN47928701, SAMN47928702, SAMN47928703, SAMN47928704, SAMN47928705, SAMN47928706, SAMN47928707, SAMN47928708, SAMN47928709, SAMN47928710. These datasets are publicly available and can be accessed through the NCBI SRA linked to each BioSample record.

提供机构:
Zenodo
创建时间:
2025-04-27
二维码
社区交流群
二维码
科研交流群
商业服务