遇见数据集

Origins of Eukaryotic Metabolism

收藏
Zenodo2026-03-26 更新2026-05-26 收录
官方服务:

资源简介:

ABSTRACT The origin of eukaryotes is a key event in the evolution of cellular life hypothesized to involve a symbiotic integration between a member of the Asgard archaea and the Alphaproteobacteria. Recent work has provided evidence for additional genetic input from other prokaryotes to the eukaryotic proteome yet the extent and sources of these contributions remain debated. Here we aimed to further resolve the prokaryotic origins of eukaryotic genes to inform our understanding of eukaryogenesis. Specifically, we developed a phylogenetic framework to investigate the origins of eukaryotic gene families associated with metabolism and informational processing for comparison. We found that informational processing genes were predominantly derived by archaea, eukaryotic metabolism is highly chimeric in its origin. In contrast to previous studies, we report a substantial number of archaeal origins of diverse metabolic enzymes including key metabolic regulators. This highlights an overlooked participation of archaeal metabolism and pinpoints potential metabolic integrations during eukaryogenesis. Apart from the alphaproteobacterial contributions to the eukaryotic metabolism, we found an additional dominant phylogenetic signal of genes potentially derived from Myxococcota, especially for gene families associated with lipid metabolism. By systematically analysing the origins of eukaryotic metabolism, this research offers novel insights into the origin of eukaryotic membranes and refine our current models for the origin of the eukaryotic cell. CONTENTS OF THE DATASET PhylogeneticTrees/ This folder contains the phylogenetic trees for the analyses. PhylogeneticTrees/CORE_dataset/ This folder contains the initial phylogenetic trees that were used to identify the potential LECA clades for metabolic (TREES_MET_INITIAL.tar.gz) and informational (TREES_INFO_INITIAL.tar.gz) set of gene families. linsi- prefix denotes those gene family phylogenies that were performed with MAFFT-linsi and strict trimming, while aln- prefix denote those gene family phylogenies that were performed with MAFFT-auto and relaxed trimming (see methods and Supplementary Figure 1). PhylogeneticTrees/EXPANDED_dataset/ This folder contains the final phylogenetic that were used to identify the sister groups of LECA clades for metabolic and informational set of gene families using empirical model (LG+G; TREES_LG*) and mixture models (LG+C20+G+F; TREES_C20*). linsi- prefix denotes those gene family phylogenies that were performed with MAFFT-linsi and strict trimming, while aln- prefix denote those gene family phylogenies that were performed with MAFFT-auto and relaxed trimming (see methods and Supplementary Figure 1). RAWDATA/ This folder contains the tables convertibles to dataframes with the data of reading programatically each individual tree. These tables can be loaded into the script to execute the analyses, instead of reading all trees. DISTANCE_PRESENCE_*, contains information for the presence of prokaryotic phyla at different topological distances (1, 1-2, 1-2-3), features of the potential LECA cades, stem branch-lengths and functional annotations (KOs). BAC_ARCH_ratio_Dist3*, contains information about the ration of bacterial/archaeal composition of the sistergroups at topological distance 3. NUMBER_OF_OGS_*, contains information about the number of monophyletic eukaryotic groups per each phylogenetic tree. BRANCHING_WITHIN_*, contains information about the prokaryotic taxa that is branching within potential LECA clades. *_ORIGINS_d123, contains the summarized information of the sister groups at different topological distance (most abundant phylum) for each LECA clade. Script/ This folder contains the jupyter notebook with the python code for analyzing the phylogenetic trees and making the respective plots (SisterGroups_at_TopologicalDistances.ipynb). Note that it requires some edits depedending of analyzing metabolic or informational protein family phylogenies (see comments). Running the analyses can take ~3 hours for metabolic set and ~5 hours for informational set. Alternatively, data can be loaded from RAWDATA/ folder with no need of reading trees. CollapsedPDFs/ This folder contains plotted phylogenetic trees (TREES_C20*). Pruned phylogenetic trees representing potential archaeal origins related with amino acid metabolisms. Tip labels of collapsed branches indicate the most abundant presence of taxa. In the same order as in labels: bacteria or archaea, total number of sequences (also indicated between brackets), number of archaeal (A-), bacterial (B-) and asgardarchaeal (Asg-) sequences, the most abundant prokaryotic phyla, and the proportion of most abundant prokaryotic phyla. For eukaryotic tip labels: KO ids, KO gene name, number of Excavata (E-), Amorphea (A-) and Diaphoretickes (D-) sequences, and total number of sequences between brackets. Only those sister groups that are at topological distance 6 to the LECA node are shown. Trees were plotted using mulberrytree (https://github.com/danieltamarit/mulberrytree).

提供机构:
Zenodo
创建时间:
2026-03-26
二维码
社区交流群
二维码
科研交流群
商业服务