Actinidia chinensis Red5 genome assembly (version 2) and annotation files
收藏资源简介:
We present version 2 of the genome assembly for <em>Actinidia chinensis</em> var. <em>chinensis</em> genotype Red5. The Red5 genome was originally assembled using short read Illumina data (Pilkington et al, 2018; https://doi.org/10.1186/s12864-018-4656-3). In version 2 we employed Pacific BioSciences Sequel Single Molecule Real Time (SMRT) sequencing technology in place of Illumina paired end read sequencing for the main assembly but leveraged that short read data (Pilkington et al, 2018) for post assembly base correction of long read assembly contigs. Additionally the Illumina long insert libraries from Pilkington et al (2018) were used for post assembly scaffolding of contigs. Scaffold assignment to linkage groups leveraged the genetic map described in Pilkington et al (2018) as well as consensus evidence from DNA synteny comparisons to existing whole genome sequences from <em>Actinidia</em>. To meet the file size restrictions some dataset components have been split into multiple parts. <strong>Assembly</strong> The assembly work flow used the FALCON/FALCON-unzip assembly suite is described in Red5_version_2_genome_assembly.md. The assembly yielded both primary and haplotig contig data sets, the metrics for which are documented in this file. The CDS and predicted peptide fasta and GFF3 gene annotation for the primary and haplotig sets are provided in separate files. <strong>File Descriptions</strong> Files named chr1.fasta to chr29.fasta represent the primary assembly linkage group level assembly units Files named haplotig_part_1.fasta to haplotig_part_10.fasta represent the haplotig contig sets split into 10 parts to meet upload file size restrictions Files named primary_assembly.primary.gff3 and haplotig.gff3 contain the gene model annotations for the primary and haplotig assembly datasets respectively primary_assembly.cds.fasta and primary_assembly.pep.fasta contain the CDS and peptide sequences for the annotations on the primary contigs haplotig.cds.fasta and haplotig.pep.fasta contain the CDS and peptide sequences for the annotations on the haplotig contigs haplotigs.placements.tsv and haplotigs.reassignments.tsv describe the placement of haplotigs relative to the primary contigs as derived from purge_haplotigs The file Red5_version_2_genome_assembly.md describes the assembly work flow and code steps used as well as assembly metrics Files HYV3_1.v.R5V2_1.png to HYV3_29.v.R5V2_29.png depict Circos plots of DNA:DNA synteny based on 1coords alignment filter of nucmer alignments using dnadiff See Red5_version_2_genome_assembly.md for description of assembly methods and assembly metrics. <strong>Funding</strong> This work was funded by Kiwifruit Royalty Investment Program by The New Zealand Institute for Plant & Food Research Ltd. with support from Zespri, and the CORE grant Endeavour Smart Idea Fund (UOOX1801) from the New Zealand Ministry of Business, Innovation and Employment (MBIE). The funding bodies had no role in the design of the study, the collection, analysis, or interpretation of data or writing this manuscript.
本研究发布中华猕猴桃(*Actinidia chinensis* var. *chinensis*)基因型Red5的第二版基因组组装结果。 Red5基因组最初采用Illumina短读长测序数据完成组装(Pilkington等,2018;https://doi.org/10.1186/s12864-018-4656-3)。第二版组装则改用太平洋生物科学(Pacific BioSciences)Sequel单分子实时测序(Single Molecule Real Time, SMRT)技术完成核心组装,而非原版本使用的Illumina双端读长测序数据;同时保留了Pilkington等(2018)发布的短读长数据,用于长读长组装重叠群(contig)的后期碱基校正。 此外,本研究还利用Pilkington等(2018)构建的Illumina大片段插入文库,完成重叠群的后期支架(scaffold)组装。将支架锚定至连锁群时,既参考了Pilkington等(2018)构建的遗传图谱,也整合了与已发布猕猴桃全基因组序列的DNA共线性比对得到的一致性证据。 为满足文件上传的大小限制,部分数据集组件被拆分为多个子文件。 **组装流程**:本研究采用FALCON/FALCON-unzip组装套件完成组装,具体流程详见Red5_version_2_genome_assembly.md文件。本次组装同时获得了核心重叠群(primary contig)和单倍型重叠群(haplotig)两套数据集,相关统计参数已在该文件中记录。针对核心重叠群和单倍型重叠群的编码序列(CDS)、预测肽段序列以及GFF3格式基因注释文件均以独立文件提供。 **文件说明**: - 命名为chr1.fasta至chr29.fasta的文件为核心组装的连锁群级组装单元; - 命名为haplotig_part_1.fasta至haplotig_part_10.fasta的文件为拆分后的单倍型重叠群数据集,共分为10个子文件以符合上传大小限制; - primary_assembly.primary.gff3与haplotig.gff3文件分别包含核心组装数据集和单倍型组装数据集的基因模型注释信息; - primary_assembly.cds.fasta与primary_assembly.pep.fasta文件包含核心重叠群注释对应的编码序列与肽段序列; - haplotig.cds.fasta与haplotig.pep.fasta文件包含单倍型重叠群注释对应的编码序列与肽段序列; - haplotigs.placements.tsv与haplotigs.reassignments.tsv文件描述了基于purge_haplotigs软件得到的单倍型重叠群相对于核心重叠群的锚定与重新分配信息; - Red5_version_2_genome_assembly.md文件详细记录了组装流程、代码步骤以及组装统计参数; - 命名为HYV3_1.v.R5V2_1.png至HYV3_29.v.R5V2_29.png的文件为Circos共线性图谱,基于dnadiff软件的nucmer比对结果,经1coords过滤后绘制得到DNA:DNA共线性可视化结果; - 关于组装方法与统计参数的详细说明,请参见Red5_version_2_genome_assembly.md文件。 **资助信息**:本研究由新西兰植物与食品研究有限公司的猕猴桃特许使用费投资计划(Kiwifruit Royalty Investment Program)资助,并得到佳沛(Zespri)的支持,同时获得新西兰商业、创新与就业部(Ministry of Business, Innovation and Employment, MBIE)的CORE专项基金“奋进智慧创意基金(UOOX1801)”资助。资助方未参与本研究的设计、数据收集、分析与解读,也未参与本手稿的撰写。



