Genome assemblies and gene annotations for the tadpole shrimp (Lepidurus arcticus)
收藏资源简介:
We provide haplotype-resolved, chromosome-level genome assemblies and gene annotations for the tadpole shrimp (Lepidurus arcticus). For each pseudo-haplotype, hap1 and hap2, we include the assembled genome sequence (FASTA) and corresponding gene annotations (GFF3). The GFF3 files are provided in their original form. For submission to the European Nucleotide Archive (ENA), these files were converted to EMBL format using EMBLmyGFF3. This conversion standardizes feature formatting and may simplify or modify some annotation details.Data generation and assembly: (https://github.com/ebp-nor/GenomeAssembly) The qbLepArct1.1 assembly is based on a combination of 144-fold coverage in Pacific Biosciences single-molecule HiFi long reads and 644-fold coverage in Arima Hi-C reads, resulting in two haplotype-separated assemblies. Assemblies were generated using hifiasm (Cheng et al. 2021) with Hi-C integration to produce two pseudo-haplotype-resolved assemblies (hap1 and hap2). Haplotype-specific assembly k-mers were identified using meryl (Rhie et al. 2020) and used to generate haplotype-filtered Hi-C read sets for scaffolding each pseudo-haplotype assembly. Hi-C reads were aligned with BWA-MEM (Li, 2013), processed with samtools (Li et al., 2009), and scaffolding was performed using YaHS (Zhou, McCarthy, and Durbin 2023). Contamination screening and removal were carried out using FCS-GX (Astashyn et al. 2024), and assemblies were manually curated using Hi-C contact maps in PretextView following the Rapid-curation-2.0 workflow. The mitochondrion was searched for in contigs and reads using Oatk (Zhou et al. 2024). Assembly quality and completeness were assessed using Merqury (Rhie et al. 2020) and BUSCO (Manni et al. 2021). Gene annotation: (https://github.com/ebp-nor/GenomeAnnotation) Genome annotation was performed using a pre-release version of the EBP-Nor annotation pipeline. Assemblies were masked for repeats using RED (Girgis, 2015) via redmask. Protein evidence was generated by extracting longest isoform sequences from Daphnia pulex using AGAT and aligning them to the assemblies with miniprot. Additional protein evidence from UniProtKB/Swiss-Prot release 2025_03 (UniProt Consortium, 2023) and the Arthropoda subset of OrthoDB v12 (Kuznetsov et al., 2023) was aligned separately. Gene prediction was performed using GALBA (Brůna et al., 2023; Buchfink et al., 2015; Hoff & Stanke, 2019; Li, 2023; Stanke et al., 2006) in miniprot mode, and Helixer was run using the invertebrate-specific model (invertebrate_v0.3_m_0100) to generate ab initio gene predictions. Protein alignment evidence and gene predictions from GALBA and Helixer were combined using EvidenceModeler via Funannotate (Haas et al., 2008). Predicted proteins were filtered against repeat-associated proteins using DIAMOND (Buchfink et al., 2015), and gene models were refined using AGAT. Functional annotation was assigned using DIAMOND searches against UniProtKB/Swiss-Prot and domain prediction with InterProScan (Jones et al., 2014). Final gene models were annotated with gene names and functional information using AGAT.



