Genome assemblies and gene annotations for the bluethroat (Luscinia svecica svecica)
收藏资源简介:
We provide haplotype-resolved, chromosome-level genome assemblies and gene annotations for the bluethroat (Luscinia svecica svecica). For each haplotype (hap1 and hap2), we include the assembled genome sequence (FASTA) and corresponding gene annotations (GFF3). The GFF3 files are provided in their original form as used for downstream analyses. For submission to the European Nucleotide Archive (ENA), files are converted to EMBL format using EMBLmyGFF3, which standardizes feature and may simplify or modify some annotation details. TE annotations (GFF3) are included as an additional resource. Data generation and assembly: (https://github.com/ebp-nor/GenomeAssembly)The bLusSve1.1 assembly is based on a combination of Oxford Nanopore Technologies (ONT) long reads (~33×) and Arima Hi-C data (~38×). ONT reads were basecalled using the Dorado SUP v.5.0.0 model, and assemblies were generated using hifiasm (Cheng et al. 2021) with Hi-C integration to produce two pseudo-haplotype-resolved assemblies (hap1 and hap2). K-mers were identified using meryl (Rhie et al. 2020) and used to partition Hi-C reads into haplotype-specific sets, ensuring haplotype-consistent scaffolding. Hi-C reads were aligned with BWA-MEM (Li, 2013), processed with samtools (Li et al., 2009), and scaffolding was performed using YaHS (Zhou, McCarthy, and Durbin 2023). Contamination screening and removal were carried out using FCS-GX (Astashyn et al. 2024), and assemblies were manually curated using using Hi-C contact maps in PretextView following the Rapid-curation-2.0 workflow. Microchromosome scaffolds were identified using MicroFinder (Mathers et al. 2025) and subjected to targeted second pass curation. Sex chromosomes and the mitochondrial genome are placed in hap1. Assembly quality and completeness were assessed using Merqury (Rhie et al. 2020) and BUSCO (Manni et al. 2021). Gene annotation: (https://github.com/ebp-nor/GenomeAnnotation)Genome annotation was performed using a pre-release version of the EBP-Nor annotation pipeline. Assemblies were masked for repeats using RED (Girgis, 2015) via redmask. Protein evidence was generated by extracting longest isoform sequences from Gallus gallus using AGAT and aligning them to the assemblies with miniprot. Additional protein evidence from UniProtKB/Swiss-Prot (UniProt Consortium, 2023) (release 2025_03) and Aves OrthoDB v12 (Kuznetsov et al., 2023) was aligned separately. Gene prediction was performed using GALBA (Brůna et al., 2023; Buchfink et al., 2015; Hoff & Stanke, 2019; Li, 2023; Stanke et al., 2006) (miniprot mode) and Helixer (Holst et al. 2025) (vertebrate model). All evidence sources were combined using EvidenceModeler via Funannotate (Haas et al., 2008). Predicted proteins were filtered against repeat-associated proteins using DIAMOND (Buchfink et al., 2015), and gene models were refined using AGAT. Functional annotation was assigned using DIAMOND searches against UniProtKB/Swiss-Prot and domain prediction with InterProScan (Jones et al., 2014). Final gene models were annotated with gene names and functional information using AGAT. MHC genes were manually curated using targeted domain-based annotation. Candidate loci were identified via Pfam domain searches with HMMER3 (Eddy 2011) on six-frame translations of each haplotype. Overlapping models were compared with Helixer predictions and known MHC genes from related species, and refined as needed. Mitochondrial genes were identified using MitoHiFi (Uliano-Silva et al. 2023). Transposable elements were annotated separately using EDTA (Ou et al. 2019).



