Supplementary materials for "First genome sequences of the dinoflagellate Effrenium voratum strain isolated from the coral Hydnophora exesa in temperate Japanese region"
收藏资源简介:
Objectives Dinoflagellates are important unicellular eukaryotes found in freshwater and marine environments. The Family Symbiodiniaceae is well-known as symbiotic algae in various marine organisms, including many endosymbiotic species living in cnidarian cells. In contrast, the genus Effrenium is non-symbiotic, exhibiting traits such as resting cyst formation and the absence of heterotrophic feeding. Effrenium strains have been used as controls in symbiosis-related comparative studies. The only described species, E. voratum, is cosmopolitan, and draft genomes of three strains have been reported. However, genomic data from strains originating in temperate Japanese regions are lacking. A high-quality draft genome of an additional E. voratum strain will provide valuable resources for molecular biology of symbiosis and comparative genomics on local adaptation in Symbiodiniaceae. Data description We describe the first draft genome (version 1) of E. voratum strain NIES-2908, which was isolated from the coral Hydnophora exesa in the temperate Japanese coastal water of Goto islands, Japan. Using PacBio HiFi sequencing, the genome was assembled to a total size of 1,006 Mbp, comprising 6,051 contigs. The GC content was 50.5%, and repetitive sequences accounted for 37.9% of the genome. Gene prediction identified 54,346 protein-coding genes, with 68.4% completeness based on the BUSCO alveolate dataset. This project contains gene models for E. voratum (NIES-2908 strain) genome assemblies annotated in this study as listed below: coding sequences (cds.fa) the longest coding sequence per gene (longest.cds.fa) protein sequences (protein.fa) the longest protein sequence per gene (longest.protein.fa) gene annotations in GTF format (.gtf) We also provide supplementary materials for the manuscript. Supplementary Figure S1 (schematic overview of the genome assembly and gene annotation workflow) Supplementary Figure S2 (genome size estimation) Supplementary Figure S3 (molecular phylogenetic tree based on 18S rRNA) Supplementary Tables (information for raw data, proteomes used for gene annotation, and genome/gene models statistics) The raw reads are available in BioProject accession PRJDB20335. The draft genome sequences are available under accession BAAHJL010000001-BAAHJL010006051. Contig.acclist.txt is the correspondence table.



