Draft genome and annotation of Raphidonema monicae
收藏资源简介:
Microalgae synthesize diverse lipids that perform a wide variety of structural, metabolic, and signalling roles in the cell. However, we know relatively little about lipid metabolism across different protist groups, and especially those adapted to extreme environments. As part of a project to investigate the lipid metabolism of Raphidonema monicae strain SAG 2030, a cold-adapted microalga isolated from Antarctica, a draft genome was assembled and annotated. The genome was assembled from a single DNA library sequenced with Illumina 150 bp paired-end reads and assembled by NovoGene (Hong Kong) with SOAPdenovo2. Novogene also perfomed curation and structural annotation of CDS and peptide sequences. The data were subsequently functionally annotated in our lab using BlastP, InterProScan and emapper.py, and the results were curated with OmicsBox software. The assembly sequence and annotation are provided as data files supporting the lipidomic investigation of the effects of abiotic stress and include the following: File 1. ‘KAD1.seq’ is the assembled genome sequence contigs that have been curated by NovoGene, .fasta format. File 2. ‘KAD1.gff’ is the annotation file .gff format for the gene CDS, .gff format. File 3. ‘KAD1.cds’ contains the CDS nucleotide sequences, .fasta format. File 4. ‘KAD1.pep’ The corresponding peptide sequences, .fasta format. File 5. ‘kad.annotation.table.xls’ is the combined functional annotation of the peptide sequences using BlastP, InterProScan and Emapper.py in an excel-readable table format.
微藻可合成多种脂质,这些脂质在细胞中承担多样的结构、代谢及信号传导功能。然而,目前学界对不同原生生物类群的脂质代谢,尤其是适应极端环境的类群的脂质代谢,所知甚少。作为研究Raphidonema monicae菌株SAG 2030脂质代谢项目的一部分,该菌株是一株分离自南极的适冷微藻,我们组装并注释了其基因组草图。 该基因组由单DNA文库构建,采用Illumina 150 bp双端测序读段进行测序,由香港诺禾致源(NovoGene)使用SOAPdenovo2完成组装。诺禾致源同时完成了编码序列(CDS)与肽序列的质控及结构注释。随后本实验室利用BLASTP、InterProScan及emapper.py对数据进行了功能注释,并通过OmicsBox软件对注释结果进行了整理质控。 本数据集提供组装序列与注释文件,用于支持非生物胁迫影响的脂质组学研究,具体文件如下: 1. 文件‘KAD1.seq’:为诺禾致源质控后的组装基因组序列重叠群,格式为FASTA(.fasta)。 2. 文件‘KAD1.gff’:为基因编码序列(CDS)的注释文件,格式为GFF(.gff)。 3. 文件‘KAD1.cds’:包含编码序列(CDS)的核苷酸序列,格式为FASTA(.fasta)。 4. 文件‘KAD1.pep’:包含对应的肽序列,格式为FASTA(.fasta)。 5. 文件‘kad.annotation.table.xls’:为肽序列经BLASTP、InterProScan及emapper.py注释后的整合功能注释表,格式为可读取的Excel表格文件。



