OKI2018_I69 assembly and annotation of the genome of an individual Oikopleura dioica from Okinawa
收藏资源简介:
A chromosome-scale assembly of the <em>Oikopleura dioica</em> genome from Okinawa, Japan. The contig assembly was generated with long-read Nanopore data using Canu pipeline v1.8, and polished with short Illumina MiSeq reads using Pilon v1.22. Both Nanopore and Illumina data were generated from DNA of a single <em>O. dioica </em>male. Hi-C chromosomal conformation capture data was used to order and orient the contigs into scaffolds using Juicer v1.6 and 3D de novo assembly (3D-DNA) pipelines. The OKI2018_I69_1.0 assembly comprises 19 scaffolds with an N50 of 16.2 Mbp (OKI2018_I69_1.0.fa). The total assembly length is 64.3 Mbp. The five longest scaffolds represent autosomal chromosomes (chr 1 and chr 2), and sex chromosomes split into pseudo-autosomal region (PAR) and X-specific (XSR) or Y-specific (YSR) regions. One of the smaller scaffolds represent a draft assembly of mitochondrial genome (chrUn_12). The rest of scaffolds are highly repetitive and were marked as unplaced. The OKI2018_I69_1.0 assembly was annotated with AUGUSTUS v3.3 and MAKER v3.01.03 pipelines. Gene predictions from these software were refined and merged using EvidenceModeler v1.1.1. To predict UTRs and alternative isoforms, the EVM models were updated using two round of PASA pipeline, resulting in 18,485 transcript models distributed among 16,936 protein-coding genes (OKI2018_I69_1.0.gene_models.gff3).
来自日本冲绳的住囊虫(Oikopleura dioica)染色体级基因组组装。本研究的重叠群(contig)组装采用纳米孔(Nanopore)长读长测序数据,通过Canu v1.8分析流程完成,并使用Illumina MiSeq短读长测序数据结合Pilon v1.22进行序列校正。纳米孔与Illumina测序数据均取自单只雄性住囊虫的基因组DNA。Hi-C染色体构象捕获数据通过Juicer v1.6与3D从头组装(3D-DNA)流程,将重叠群排序并定向至支架(scaffold)序列。本次组装版本OKI2018_I69_1.0包含19条支架序列,N50值为16.2 Mbp(对应组装文件为OKI2018_I69_1.0.fa),总组装长度为64.3 Mbp。其中最长的5条支架序列为常染色体(chr 1与chr 2),性染色体则被划分为假常染色体区域(PAR)、X染色体特异区域(XSR)与Y染色体特异区域(YSR)。较小的支架序列之一为线粒体基因组草图组装(chrUn_12),剩余支架序列均为高度重复序列,被标记为未定位序列。OKI2018_I69_1.0组装版本通过AUGUSTUS v3.3与MAKER v3.01.03流程完成基因注释。通过EvidenceModeler v1.1.1(EVM)对上述两款软件预测的基因结果进行优化与整合。为预测非翻译区(UTR)与可变剪接异构体,采用两轮PASA分析流程更新EVM模型,最终得到分布于16936个蛋白编码基因上的18485个转录本模型(对应注释文件为OKI2018_I69_1.0.gene_models.gff3)。



