The First Highly Contiguous Genome Assembly of Pikeperch (Sander lucioperca), an Emerging Aquaculture Species in Europe
收藏资源简介:
<strong>Supporting data for "The First Highly Contiguous Genome Assembly of Pikeperch (<em>Sander lucioperca</em>), an Emerging Aquaculture Species in Europe"</strong> =========================================================================================== <strong>Abstract:</strong> -------- The pikeperch (<em>Sander lucioperca</em>) is a fresh and brackish water Percid fish natively inhabiting the northern hemisphere. This species is emerging as a promising candidate for intensive aquaculture production in Europe. Specific traits like cannibalism, growth rate and meat quality require genomics based understanding, for an optimal husbandry and domestication process. Still, the aquaculture community is lacking an annotated genome sequence to facilitate genome-wide studies on pikeperch. Here, we report the first highly contiguous draft genome assembly <em>S. lucioperca</em>. In total, 413 and 66 giga base pairs of DNA sequencing raw data were generated with Illumina platform and PacBio Sequel System, respectively. The PacBio data were assembled into a final assembly size of ~900 Mb covering 89% of the 1,014 Mb estimated genome size. The draft genome consisted of 1,966 contigs ordered into 1,313 scaffolds. The contig and scaffold N50 lengths are 3.0 Mb and 4.9 Mb, respectively. The identified repetitive structures accounted for 39% of the genome. We utilized homologies to other ray-finned fishes, and ab initio gene prediction methods to predict 21,249 protein-coding genes in the <em>S. lucioperca </em>genome, of which 88% were functionally annotated by either sequence homology or protein domains and signatures search. The assembled genome spans 97.6% and 96.3% of Vertebrate respectively Actinopterygii single-copy orthologs. The outstanding mapping rate (99.9%) of genomic PE-reads on the assembly suggests an accurate and nearly complete genome reconstruction. This draft genome sequence is the first genomic resource for this promising aquaculture species. It will provide an impetus for genomic-based breeding studies targeting phenotypic and performance traits of captive pikeperch. <strong>Files:</strong> ------ sanlu.cds.renamed.fa - Coding sequences of predicted protein-coding genes sanlu.genes.filt.gff3 - gff3 file of predicted protein coding genes sanlu.genes.pep.fa - predicted peptide sequences sanlu.genome.ctg.fasta - <em>Sander lucioperca</em> genome assembly at contig-level sanlu.genome.scf.fa - <em>Sander lucioperca</em> genome assembly at scaffold-level Sanlu.genome.masked.fasta - Repeats-masked <em>Sander lucioperca</em> genome assembly at scaffold-level Sanlu.genome.repeats.gff - Gff3 file of predicted repeats in <em>Sander lucioperca</em> genome Additional_File_2.xlsx - Functional annotations of <em>Sander lucioperca </em>genes by SwissProt, NR RefSeq, TrEMBL and InterPro databases sanlu.repeats.lib.fasta - Predicted repeats library in <em>Sander lucioperca </em>in FASTA format sanlu_miRNA.csv Predicted micro RNA families in CSV tab file sanlu_miRNA.bed - Predicted micro RNA families in BED file format sanlu_miRNA.html - Predicted micro RNA families in HTML sanlu_rRNA.fasta - Predicted ribosomal RNA (rRNA) sequences in FASTA file format sanlu_rRNA.gff - Predicted ribosomal RNA (rRNA) sequences in GFF file format trna.genes.csv - Predicted transfer RNA (tRNA) genes in CSV tab file SpeciesTree_rooted_node_labels.txt - Predicted phylogenetic tree in NEWICK format SpeciesTreeAlignment.fa - Species tree alignment in FASTA, based on 1.1 single copy orthologs



