遇见数据集

New genomic insights into long-chain alkenone biosynthesis by the coccolithophorid marine alga Gephyrocapsa huxleyi

收藏
Zenodo2026-06-13 更新2026-06-17 收录
官方服务:

资源简介:

Abstract Background Gephyrocapsa huxleyi is a coccolithophoric haptophyte widespread across most marine ecosystems that plays an important role in the carbon cycle. G. huxleyi is one of the few haptophytes known to produce long-chain alkenones (LCAs), which are lipids composed of C35-C42 n-alkyl chains, with two to four trans-double bonds and a keto group at either the 2nd or 3rd carbon positions, which are widely used as proxies for paleotemperature reconstruction. Despite the biomarker value of LCAs, little is known about their biosynthesis, which severely limits the understanding of the regulation of their production under different conditions, and thus their predictive nature. Differences in gene expression under different conditions may help to unravel the LCA biosynthetic pathway, but require annotated reference genomes. Although a large number of G. huxleyi strains inhabiting diverse ecosystems exist, only one such genome (i.e. of strain CCMP1516) is currently available. Results To evaluate differences in LCA biosynthesis under changing conditions, we developed a tailored approach to decontaminate and assemble the genomes of two additional non-axenic G. huxleyi strains (strains CCMP1742 and CCMP2758) that produce LCAs with two to four double bonds, one of which (CCMP2758) also produces unusually short (C35-C36) LCAs. We identified a polyketide synthase (PKS) in strain CCMP1516, which we propose as a candidate to carry out LCA formation based on the structure and composition of its modules. Additionally, we identified several CCMP1742 and CCMP2758 PKSs that were differentially expressed between temperatures and across time that closely resemble this PKS and are likely to be involved in LCA formation. Based on differences in gene expression between two different growth temperatures, we also identified a limited number of desaturases which are potentially responsible for the addition of the third and fourth double bonds in the LCAs produced by these strains. Conclusions Here, we advance the understanding of biosynthesis of LCAs, important molecules for paleotemperature reconstruction, in the marine haptophyte G. huxleyi. By assembling the genomes of two LCA-producing strains, we were able to overcome the limitations of the scarcity of reference genomes of this group, enabling the analysis of gene expression under varying growth conditions and the identification of PKSs and desaturases potentially involved in LCA biosynthesis These findings constitute a step forward in elucidating the genetic basis for LCA production and regulation in G. huxleyi, which not only advances our understanding on how LCAs are formed but also aids in the improvement of LCA-based paleotemperature reconstructions. Here we supply: Supplementary Data 1. Gas Chromatography (GC) and GC-Mass Spectrometry (GC-MS) lipid data files. This folder contains raw data for all lipid analyses. The files are organized into subfolders based on the figure they support. The dataset includes the raw GC files used for quantification and the GC-MS files used for compound identification, including the specific chromatograms that generated the representative peaks shown in Figure 2. Where applicable, the subfolders also contains a summary table with the quantitative results extracted from the chromatograms and a sample metadata file that lists the experimental conditions (e.g., temperature, time point) and links them to the corresponding data file. Supplementary Data 2. De novo genome assemblies of G. huxleyi strains CCMP1742 and CCMP2758 generated during the evaluation of various assembly tools shown in Supplementary Figure 2. These include long read assemblers (CANU v1.8, flye v2.8.1, shasta v0.5.1 and wtdbg2 v2.3 with the wtpoa-cns v2.3 consenser), the short read assembler SPAdes v3.14.1, and hybrid assemblers that use both short and long reads (Wengan v0.2—using either DiscovarDeNovo [WenganD] or Minia3 [WenganM] as the short-read assembler—hybridSPAdes v3.14.1, biosyntheticSPAdes v 3.14.1, and HASLR v 0.8a1. and a metagenomic assembler (OPERA-MS (v 0.8.2) in hybrid mode using MegaHIT as short read assembler. A metagenomic assembler, OPERA-MS v0.8.2 run in hybrid mode with MegaHIT, was also evaluated. Furthermore, two scaffolders, WenganD and OPERA-LG v2.0.5, were applied to two selected assemblies (hybridSPAdes and biosyntheticSPAdes). Supplementary Data 3. De novo genome and metagenome assemblies of G. huxleyi strains CCMP1742 and CCMP2758 generated at different stages of the assembly process outlined in Figure 3. The included assemblies correspond to specific steps: (1) the initial hybrid metagenomic assemblies generated using MetaSPAdes (v 3.14.1) and OPERA-MS (with SPAdes as the short-read assembler); (3) the hybrid assemblies created with SPAdes and biosyntheticSPAdes after a read-filtering step, where reads were mapped to the metagenomic assembly using Minimap2 and those mapping to contigs identified as contamination by CAT classification, coverage, and GC content were removed; (4) the scaffolded assemblies generated using OPERA-LG; and (7) the final assemblies after the identification and removal of remaining non-eukaryotic scaffolds. Supplementary Data 4. Funannotate genome annotation output files. This folder contains the “annotate_results” directories generated by Funannotate for the CCMP1742 and CCMP2758 assemblies. These assemblies were produced using MetaSPAdes as a hybrid metagenomic assembler and biosyntheticSPAdes as a hybrid genome assembler, from which reads and contigs identified as likely contamination have been removed as described in Figure 3. Supplementary Data 5. Cell count data for cold-shock and cold-adaptation experiments in G. huxleyi strains CCMP1742 and CCMP2758. This folder contains three files: two with data from the cold-shock experiments (one for each strain, CCMP1742 and CCMP2758) and one with data from the temperature acclimation experiment. The cold-shock datasets provide a summary of cell counts and viability measurements. For each sample, cell abundance was first measured on an unstained aliquot based on chlorophyll red autofluorescence (PerCP) versus forward scatter (FSC). A separate aliquot was then stained with SYTOX Green nucleic acid stain to determine cell viability, with a formalin-killed control used to define the gating strategy for dead cells. The data includes results from both untreated samples (“FCM” tab) and live/dead tests (“FCM_Live_dead” tab). Technical replicates for these experiments are included where applicable. The temperature acclimation dataset describes the cellular abundance data for strains grown at 20 °C and 7 °C. It reports raw counts for two distinct cell clusters ("Up" and "Down"), identified by their chlorophyll red autofluorescence (PerCP) versus forward scatter (FSC), and the total cell sum of both clusters for each sample. Data are presented as individual technical replicates ("raw" tabs), the average and standard deviation of those replicates per flask ("per_flask" tabs), and the average and standard deviation for each time point across three biological replicates (except t=0, n=1). Supplementary Data 6. Source data and statistical analysis for Figures 5, 6, and 7, and Supplementary Figures 10, 15, 16, 23, 25 and 31. This file contains the underlying dataset and the results of the statistical analyses used to generate the indicated figures. The dataset includes the underlying numerical values and categorical labels plotted in the figures, as well as the full outputs of statistical analyses—including test statistics, degrees of freedom, p-values, confidence intervals, and post-hoc comparisons where applicable. Supplementary Data 7. RNA-seq alignment and deduplication metrics. This file contains the output metrics from the STAR aligner and UMI-tools deduplication for RNA-seq data. It includes key alignment statistics—such as the number of input reads, uniquely mapped reads (%), and reads mapped to multiple loci (%)—for strains CCMP1742 and CCMP2758 before and after deduplication, and for strain CCMP1516 (non-deduplicated). The file also includes the deduplication results from running “umi_tools dedup” command for samples of CCMP1742 and CCMP2758 strains, reporting metrics such as number of input reads, final output counts and mean and maximum number of unique UMIs per position. Supplementary Data 8. Phylogeny of desaturases. This folder contains the file used to generate Figure 9 and Supplementary Figure 33. The folder contains the alignment and trimmed alignment, IQ-TREE output files, and iTOL annotation files.

提供机构:
Zenodo
创建时间:
2026-06-13
二维码
社区交流群
二维码
科研交流群
商业服务