Genome assembly and annotation of the 13-year periodical cicada Magicicada neotredecim (Insecta: Hemiptera: Cicadidae)
收藏资源简介:
We assembled the genome of a 13-year periodical cicada Magicicada neotredecim as follows. Total genomic DNA was extracted from a male of M. neotredecim (Brood XIX) collected from Illinois in 2011. Approximately 800 Gb of genomic sequence data were generated from a 10x Genomics Chromium linked-read library using the Illumina NovaSeq 6000 platform. We constructed draft contigs as follows. Because the 10x Genomics linked-read dataset was sequenced at very high depth, barcode-aware de novo assemblers were not computationally tractable within the available memory limits. We therefore generated an initial draft assembly using MEGAHIT v1.2.9 (Li et al., 2015) on the full Illumina read set while ignoring barcode information, producing a contig set suitable for downstream long-range scaffolding. To recover long-range linkage information from the 10x barcodes, linked reads were mapped back to the MEGAHIT contigs, and barcode molecule information was extracted. Putative misassemblies were identified and broken using Tigmint v1.2.10 (Jackman et al., 2018) following the standard tigmint-make workflow, which detected contigs with inconsistent molecule support and split at unsupported breakpoints. The corrected contigs were then scaffolded using barcode evidence with ARCS v1.1.0 (Yeo et al., 2018). Scaffolding was performed conservatively to prioritize joins supported by consistent molecule linkage and to minimize joins driven by repetitive sequences. We next performed a reference-guided ordering into linkage groups. Barcode-scaffolded sequences were organized into chromosome-scale linkage groups using RagTag2 v2.1.0 (Alonge et al., 2022), with the Magicicada septendecim chromosome-level assembly (doi: 10.5281/zenodo.18828033) as a reference. RagTag was run in scaffold mode with minimap2 v2.24 (Li, 2018) using assembly-to-assembly alignment parameters (-x asm20). This step provided an initial ordering and orientation of scaffolds into putative linkage groups based on conserved synteny. To resolve scaffold gaps without introducing long-read–style gap-filling assumptions, we used a previously generated linked-read assembly produced with Supernova v2.1.1 (Weisenfeld et al., 2017) as an independent source of contiguous sequence. Supernova scaffolds were treated as an alternative assembly and used specifically to patch gaps in the RagTag-ordered assembly. Gap patching was performed with RagTag2 in patch mode, using the RagTag-ordered assembly as the target and the Supernova scaffolds as the patch source. Patches were applied conservatively: insertions were accepted only when supported by consistent, high-confidence alignments of Supernova scaffolds to both flanks of a target gap. This approach preferentially replaces placeholder N-runs with supported sequence where bridging evidence exists, while leaving unresolved gaps intact. For gene annotation, annotations from M. septendecim were transferred to the final chromosome-anchored M. neotredecim assembly using Lifton v1.0.7 (Chao et al., 2025). The curated M. septendecim GFF annotation set was used as input, and copy-aware mapping (-copies) was enabled to retain duplicated gene models when supported. Predicted genes were functionally annotated by sequence similarity searches against Drosophila melanogaster and Hemiptera protein databases using DIAMOND v2.1.9 (Buchfink et al., 2021).



