遇见数据集

The First Highly Contiguous Genome Assembly of Pikeperch (Sander lucioperca), an Emerging Aquaculture Species in Europe

收藏
Zenodo2020-07-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>Supporting data for "The First Highly Contiguous Genome Assembly of Pikeperch (<em>Sander lucioperca</em>), an Emerging Aquaculture Species in Europe"</strong> =========================================================================================== <strong>Abstract:</strong> -------- The pikeperch (<em>Sander lucioperca</em>) is a fresh and brackish water Percid fish natively inhabiting the northern hemisphere. This species is emerging as a promising candidate for intensive aquaculture production in Europe. Specific traits like cannibalism, growth rate and meat quality require genomics based understanding, for an optimal husbandry and domestication process. Still, the aquaculture community is lacking an annotated genome sequence to facilitate genome-wide studies on pikeperch. Here, we report the first highly contiguous draft genome assembly <em>S. lucioperca</em>. In total, 413 and 66 giga base pairs of DNA sequencing raw data were generated with Illumina platform and PacBio Sequel System, respectively. The PacBio data were assembled into a final assembly size of ~900 Mb covering 89% of the 1,014 Mb estimated genome size. The draft genome consisted of 1,966 contigs ordered into 1,313 scaffolds. The contig and scaffold N50 lengths are 3.0 Mb and 4.9 Mb, respectively. The identified repetitive structures accounted for 39% of the genome. We utilized homologies to other ray-finned fishes, and ab initio gene prediction methods to predict 21,249 protein-coding genes in the <em>S. lucioperca </em>genome, of which 88% were functionally annotated by either sequence homology or protein domains and signatures search. The assembled genome spans 97.6% and 96.3% of Vertebrate respectively Actinopterygii single-copy orthologs. The outstanding mapping rate (99.9%) of genomic PE-reads on the assembly suggests an accurate and nearly complete genome reconstruction. This draft genome sequence is the first genomic resource for this promising aquaculture species. It will provide an impetus for genomic-based breeding studies targeting phenotypic and performance traits of captive pikeperch. <strong>Files:</strong> ------ sanlu.cds.renamed.fa - Coding sequences of predicted protein-coding genes sanlu.genes.filt.gff3 - gff3 file of predicted protein coding genes sanlu.genes.pep.fa - predicted peptide sequences sanlu.genome.ctg.fasta - <em>Sander lucioperca</em> genome assembly at contig-level sanlu.genome.scf.fa - <em>Sander lucioperca</em> genome assembly at scaffold-level Sanlu.genome.masked.fasta - Repeats-masked <em>Sander lucioperca</em> genome assembly at scaffold-level Sanlu.genome.repeats.gff - Gff3 file of predicted repeats in <em>Sander lucioperca</em> genome Additional_File_2.xlsx - Functional annotations of <em>Sander lucioperca </em>genes by SwissProt, NR RefSeq, TrEMBL and InterPro databases sanlu.repeats.lib.fasta - Predicted repeats library in <em>Sander lucioperca </em>in FASTA format sanlu_miRNA.csv Predicted micro RNA families in CSV tab file sanlu_miRNA.bed - Predicted micro RNA families in BED file format sanlu_miRNA.html - Predicted micro RNA families in HTML sanlu_rRNA.fasta - Predicted ribosomal RNA (rRNA) sequences in FASTA file format sanlu_rRNA.gff - Predicted ribosomal RNA (rRNA) sequences in GFF file format trna.genes.csv - Predicted transfer RNA (tRNA) genes in CSV tab file SpeciesTree_rooted_node_labels.txt - Predicted phylogenetic tree in NEWICK format SpeciesTreeAlignment.fa - Species tree alignment in FASTA, based on 1.1 single copy orthologs

《欧洲新兴水产养殖物种梭鲈(<em>Sander lucioperca</em>)首个高连续性基因组组装配套支持数据》 ========================================================================================== <strong>摘要:</strong> -------- 梭鲈(<em>Sander lucioperca</em>)是一种栖息于北半球的淡水及咸水性鲈科鱼类,该物种正逐步成为欧洲集约化水产养殖极具潜力的候选品种。其同类相食、生长速率与肉质品质等特异性状,亟需依托基因组学手段开展解析,以优化养殖管理与驯化流程。然而当前水产学界仍缺乏可供梭鲈全基因组研究使用的注释版基因组序列。本研究首次报道了<em>S. lucioperca</em>的高连续性草图基因组组装。本研究分别通过Illumina测序平台与PacBio Sequel系统,共计生成413 Gb与66 Gb的DNA测序原始数据。将PacBio数据进行组装后,最终组装基因组大小约为900 Mb,覆盖了预估基因组大小(1014 Mb)的89%。该草图基因组包含1966个重叠群(contig),并被锚定至1313个支架(scaffold)。重叠群与支架的N50长度分别为3.0 Mb与4.9 Mb。鉴定得到的重复序列占基因组总长的39%。本研究借助与其他辐鳍鱼类的同源序列比对,结合从头基因预测方法,在<em>S. lucioperca</em>基因组中预测得到21249个蛋白质编码基因,其中88%的基因可通过序列同源性比对或蛋白质结构域与特征搜索获得功能注释。组装完成的基因组覆盖了97.6%的脊椎动物(Vertebrate)单拷贝同源基因以及96.3%的辐鳍鱼纲(Actinopterygii)单拷贝同源基因。将基因组PE测序reads比对至该组装序列的比对率高达99.9%,证明本研究实现了高精度且近乎完整的基因组重构。该草图基因组序列是这一极具潜力的水产养殖物种的首个基因组资源,将为针对养殖梭鲈的表型与生产性状开展基因组辅助育种研究提供重要支撑。 <strong>数据集文件列表:</strong> ------ - `sanlu.cds.renamed.fa`:预测蛋白质编码基因的编码序列文件 - `sanlu.genes.filt.gff3`:过滤后的预测蛋白质编码基因的GFF3格式文件 - `sanlu.genes.pep.fa`:预测肽序列文件 - `sanlu.genome.ctg.fasta`:<em>Sander lucioperca</em> 重叠群水平基因组组装FASTA文件 - `sanlu.genome.scf.fa`:<em>Sander lucioperca</em> 支架水平基因组组装FASTA文件 - `Sanlu.genome.masked.fasta`:经重复序列屏蔽的<em>Sander lucioperca</em> 支架水平基因组组装FASTA文件 - `Sanlu.genome.repeats.gff`:<em>Sander lucioperca</em> 基因组预测重复序列的GFF3格式文件 - `Additional_File_2.xlsx`:通过SwissProt、NR RefSeq、TrEMBL及InterPro数据库注释得到的<em>Sander lucioperca</em> 基因功能注释文件 - `sanlu.repeats.lib.fasta`:FASTA格式的<em>Sander lucioperca</em> 预测重复序列文库文件 - `sanlu_miRNA.csv`:CSV格式的预测微小RNA(miRNA)家族文件 - `sanlu_miRNA.bed`:BED格式的预测微小RNA(miRNA)家族文件 - `sanlu_miRNA.html`:HTML格式的预测微小RNA(miRNA)家族结果文件 - `sanlu_rRNA.fasta`:FASTA格式的预测核糖体RNA(rRNA)序列文件 - `sanlu_rRNA.gff`:GFF格式的预测核糖体RNA(rRNA)序列文件 - `trna.genes.csv`:CSV格式的预测转运RNA(tRNA)基因文件 - `SpeciesTree_rooted_node_labels.txt`:NEWICK格式的预测系统发育树文件 - `SpeciesTreeAlignment.fa`:基于1.1套单拷贝同源基因构建的系统发育树比对序列FASTA文件

提供机构:
Zenodo
创建时间:
2019-07-22
二维码
社区交流群
二维码
科研交流群
商业服务