Inferring and comparing metabolisms across heterogeneous sets of annotated genomes using AuCoMe
收藏资源简介:
CONTENT OF THIS ARCHIVE The Zenodo archive is composed of one file and four main directories:<br> * <strong>analyses</strong> gathers three subdirectories: algae, bacteria, and fungi. It includes all files used to create the figures, supplemental figures, and results of the paper. * <strong>code</strong> contains all AuCoMe and PADMET codes. – <strong>aucome_v0.5.1</strong> this directory gathers the code of AuCoMe used to run the three datasets. – <strong>padmet_v5.0.1</strong> this directory contains the code of PADMET used to run AuCoMe. * <strong>datasets</strong> this directory gathers all datasets on which AuCoMe was run: the bacterial, fungal, and algal datasets, and the 32 synthetic datasets, which contain an <em>E. coli</em> K–12 MG1655 genome to which various degradations were applied, together with 28 other bacterial genomes. It also encompasses the version 23.5 of MetaCyc database. * <strong>scripts_analyses</strong> this directory contains several scripts to generate the figures, supplemental figures and a script to degrade the <em>E. coli</em> K–12 MG1655 genome. 1/ Content of the <strong>analyses</strong> repertory<br> It is composed of three subdirectories: <strong>algae</strong>, <strong>bacteria</strong>, and <strong>fungi</strong>. 1.1/ Content of the <strong>algae</strong> subdirectory<br> It encompasses 9 files. * <strong>Figure_2_algal_nb_reactions.tsv</strong> for each species of the algal dataset, this file gives the number of reactions at each AuCoMe step. It was used to create figure 2D. * <strong>Figure_S10_Deepec_algal.tsv</strong> for each species of the algal dataset, at each AuCoMe step (robust orthology, non-robust orthology, and annotation or orthology), several measures were computed, i.e.: the number of reactions, the number of ECs, the number of ECs validated by DeepEC, and the ratio number of ECs valided by DeepEC / number of ECs. It was used to design figure S10(b). * <strong>Table_S6_50_random_reactions_found.xlsx</strong> contains manual validation of 50 randomly chosen reactions found in any of the species (is the Supplemental Table S6). * <strong>Table_S7_50_random reactions absent.xlsx</strong> includes manual validation of 50 reactions absent from a species and randomly chosen (is the Supplemental Table S7). * <strong>Table_S8_reactions_common_only_Cokamuranus_Sjaponica.xlsx</strong> encompasses reactions common to <em>Saccharina japonica</em> and <em>Cladosiphon okamuranus</em> but not found in other brown algae (is the Supplemental Table S8). * <strong>Table_S9_homologues_Esiliculosus_Sjaponica.xlsx</strong> contains additional homologs in <em>E. siliculosus</em> found by BLASTP searches for sequences inferred to be present only in <em>C. okamuranus</em> and <em>S. japonica</em> (is the Supplemental Table S9). * <strong>Table_S10_o-aminophenol_Esiliculosus_holomogues.xlsx</strong> includes additional o-aminophenol oxidases from <em>E. siliculosus</em> and their homologs in other stramenopiles. It is the Supplemental Table S10 with more detail (like the amino acid sequences). * <strong>Table_S11_reactions_cryptophytes_haptophytes_stramenopiles_archeplastida.xlsx</strong> encompasses reactions distinguishing the cryptophyte, haptophyte, stramenopile, and archeplastida groups (is the Supplemental Table S11). * <strong>Table_S12_pathways_cryptophytes_haptophytes_stramenopiles_archeplastida.xlsx</strong> contains shared metabolic pathways as well as the absence of pathways between chryptophytes, haptophytes, stramenopiles, and archaeplastida (is the Supplemental Table S12). 1.2/ Content of the <strong>bacteria</strong> subdirectory<br> It gathers 12 files and 9 repertories. * <strong>aucome_final.tsv</strong> output file of the figure S4 comparison bacteria.py script, for each of the 29 bacterial metabolic networks produced with AuCoMe, this table contains the number of ECs, the number of unique ECs, the number of total reactions, the number of enzymatic reactions with genes, the number of enzymatic reactions without genes, and the number of spontaneous reactions. * <strong>carveme_stat.tsv</strong> output file of the figure S4 comparison bacteria.py script, for each of the 29 bacterial metabolic networks produced with CarveMe, this table contains the number of ECs, the number of unique ECs, the number of total reactions, the number of enzymatic reactions with genes, the number of enzymatic reactions without genes, and the number of spontaneous reactions. * <strong>ecocyc.padmet</strong> contains the EcoCyc database version 23.5 at the PADMet, is used to generate the Supplemental Fig. S5. * <strong>Figure_2_bacterial_nb_reactions.tsv</strong> for each species of the bacterial dataset, this file gives the number of reactions at each AuCoMe step. It was used to create figure 2B. * <strong>Figure_3_nb_reactions_step.tsv</strong> for each dataset of the 32 synthetic bacterial datasets, this file enumerates the number of reactions at each AuCoMe step. It was used to create figure 3A. * <strong>Figure_3_fmeasure_steps.tsv</strong> for each dataset of the 32 synthetic bacterial datasets, this file indicates the values of the F-measures resulting of the comparison of the GSMNs recovered for each <em>E. coli</em> K–12 MG1655 genome replicate with the gold-standard network EcoCyc. It was used to create figure 3B. * <strong>Figure_S4_output</strong> contains 3 output files of the figure_S4_comparison_bacteria.py script: – <strong>Figure_S4_boxplot_networks.svg</strong> is the Supplemental Figure S4 in high resolution. – <strong>Figure_S4_boxplot_networks.tsv</strong> contains the number of reactions, the type of reactions (All, Reactions with genes, ...), and the used software. Thes data were produced and used in the figure_S4_comparison_bacteria.py script. – <strong>Figure_S4_barplot_time_networks.svg</strong> for each software, shows the required time in seconds used to reconstruct these bacterial metabolic networks. * <strong>Figure_S5_output</strong> encompasses 3 output files of figure S5 reference catalog.py script: – <strong>Figure_S5_ec_union.svg</strong> is the Supplemental Figure S5 in high resolution. – <strong>Figure_S5_ec_union_venn.svg</strong> another visualisation of presenting the results of the Supplemental Fig. S5. – <strong>Figure_S5_refence_ec_catalog_K12MG1655.tsv</strong> contains an EC catalog to <em>E. coli</em> K-12 MG1655 from the BIGG, EcoCyc, KEGG, and ModelSEED databases. This file is used to produce the Supplemental Figure S5. * <strong>Figure_S6_output</strong> includes 2 output files of the figure_S6.py script: – <strong>Figure_S6_comparison_all.svg</strong> is the Supplemental Figure S6 in high resolution. – <strong>Figure_S6_comparison_all.tsv</strong> contains data used to produce the Supplemental Figure S6. * <strong>gapseq_stat.tsv</strong> output file of the figure_S4_comparison_bacteria.py script, for each of the 29 bacterial metabolic networks produced with gapseq, this table contains the number of ECs, the number of unique ECs, the number of total reactions, the number of enzymatic reactions with genes, the number of enzymatic reactions without genes, and the number of spontaneous reactions. * <strong>jsons_bigg</strong> todate contains the five metabolic networks of <em>E. coli</em> K–12 MG1655 that can find in BIGG at JSON format. These files correspond to the BIGG reference metabolic network on the Supplemental Figure S5. * <strong>jsons_modelseed</strong> todate includes the metabolic network of <em>E. coli</em> K–12 MG1655 that can find in ModelSEED at JSON format. It is the ModelSEED reference metabolic network on the Supplemental Figure S5. * <strong>kegg_ecs.txt</strong> input file of the figure_S5_reference_catalog.py script, it contains matches between EC numbers and all the entries of <em>E. coli</em> K–12 MG1655 in the KEGG database. * <strong>mapping_modelseed_ec.tsv</strong> input file of the figure_S4_comparison_bacteria.py script, it encompasses matches between ModelSEED reactions and EC numbers. * <strong>modelseed_stat.tsv</strong> output file of the figure_S4_comparison_bacteria.py script, for each of the 29 bacterial metabolic networks produced with ModelSEED, this table contains the number of ECs, the number of unique ECs, the number of total reactions, the number of enzymatic reactions with genes, the number of enzymatic reactions without genes, and the number of spontaneous reactions. * <strong>networks_aucome</strong> for each of the 29 bacteria, contains a metabolic networks at the PADMet format obtained with AuCoMe. * <strong>networks_carveme</strong> for each of the 29 bacteria, contains a metabolic networks at the SBML format got to CarveMe. * <strong>networks_gapseq</strong> composes of 29 subdirectories (one for each bacterium). All these subdirectories contain 10 files about the metabolic networks a obtained with gapseq: – <strong>species-all-Pathways.tbl</strong> encompasses data on pathways at TBL format. – <strong>species-all-Reactions.tbl</strong> includes data on reactions at TBL format. – <strong>species-draft.RDS</strong> is a draft metabolic network at RDS (R Data Format). – <strong>species-draft.xml</strong> is a draft metabolic network at SBML format. – <strong>species-medium.csv</strong> encompasses all the metabolites allow the default medium. – <strong>species.RDS</strong> is the final metabolic network at RDS (R Data Format). – <strong>species-rxnWeights.RDS</strong> is a temporary file nedeed to gapseq fill at RDS (R Data Format). – <strong>species-rxnXgenes.RDS</strong> is a temporary file nedeed to gapseq fill at RDS (R Data Format). – <strong>species-Transporter.tbl</strong> includes data on transporters at TBL format. – <strong>species.xml</strong> is the final metabolic network at SBML format. * <strong>networks_modelseed</strong> includes two subdirectories: – <strong>sbml</strong> for each of the 29 bacteria, encompasses a metabolic networks at the SBML format got to ModelSEED. – <strong>tsv</strong> for each of the 29 bacteria, contains two TSV files: – <strong>genomeset__species.gbk_genome.fbamodel-compounds.tsv</strong> includes data on compounds at TSV format. – <strong>genomeset__species.gbk_genome.fbamodel-reactions.tsv</strong> encompasses data on reactions at TSV format. * <strong>time_carveme.txt</strong> input file of the figure S4 comparison bacteria.py script, for each of the 29 bacteria it stores the running time of CarveMe (in seconds) to reconstruct a metabolic network. * <strong>time_gapseq.txt</strong> input file of the figure S4 comparison bacteria.py script, for each of the 29 bacteria it stores the running time of gapseq (in seconds) to reconstruct a metabolic network. 1.3/ Content of the <strong>fungi</strong> repertory<br> It contains three files and five directories. * <strong>All-pathways-of-S.-cerevisiae-S288c.txt</strong> encompasses all the YeastCyc pathways. * <strong>Figure_2_fungal_nb_reactions.tsv</strong> for each species of the fungal dataset, this file gives the number of reactions at each AuCoMe step. It was used to create figure 2C. * <strong>Figure_S7_output</strong> contains 11 output files of the figure S7 comparison pathway fungi.py script: – <strong>completion_pathway_species.svg</strong> for each of the 5 fungi (<em>L. bicolor</em>, <em>N. crassa</em>, <em>R. oryzae</em>, <em>S. cerevisiae</em> S288C, and <em>S. pombe</em>), contains a subfigure of the Supplemental Fig. S7. – <strong>fungi_stats.tsv</strong> is the Supplemental Table S5. – <strong>pathway_venn_species.png</strong> for each of the 5 fungi (<em>L. bicolor</em>, <em>N. crassa</em>, <em>R. oryzae</em>, <em>S. cerevisiae</em> S288C, and <em>S. pombe</em>), includes a Venn diagram about all the pathways found with the 3 software (AuCoMe, gapseq, and ModelSEED). * <strong>Figures_S8_S9_output</strong> contains 11 files, in all these files, a comparison of all pathways of metabolic networks of <em>S. cerevisiae</em> S288C obtained with AuCoMe and gapseq to those of YeastCyc was released. – <strong>comparison_yeastcyc.png</strong> is a picture about number of pathways true positive, false positive, and false negative are found, according the used method (AuCoMe and gapseq). – <strong>completion_pathway_gapseq.svg</strong> includes the number of pathways common or specific to YeastCyc and gapseq with their completeness ratio predicted by gapseq. – <strong>Figure_S8_completion_pathway_aucome.svg</strong> contanis the number of pathways common or specific to YeastCyc and AuCoMe with their completeness ratio predicted by AuCoMe, is the Supplemental Figure S8. – <strong>Figure_S9_venn_diagram_70_100.svg</strong> is the Supplemental Figure S9. All pathways of AuCoMe, gapseq and YeastCyc with a completion rate between 50% and 70% are compared. – <strong>venn_diagram.svg</strong> in this picture, all pathways are compared. – <strong>venn_diagram 50.svg</strong> all pathways of AuCoMe, gapseq and YeastCyc with a completion rate less than 50% are compared. – <strong>venn_diagram_50_gapseq.svg</strong> all pathways of gapseq whatever their completion rate are compared to the AuCoMe and YeastCyc pathways with a completion rate less than 50%. – <strong>venn_diagram_50_70.svg</strong> all pathways of AuCoMe, gapseq, and YeastCyc with a completion rate between 50% and 70% are compared. – <strong>venn_diagram_50_70_gapseq.svg</strong> all pathways of gapseq whatever their completion rate are compared to the AuCoMe and YeastCyc pathways with a completion rate between 50% and 70%. – <strong>venn diagram_70_100_gapseq.svg</strong> all pathways of gapseq whatever their completion rate are compared to the AuCoMe and YeastCyc pathways with a completion rate between 70% and 100%. – <strong>yeast_cyc_comparison.tsv</strong> contains the number of pathways true positive, false positive, and false negative are found, according the used method (AuCoMe and gapseq). * <strong>Figure_S10_Deepec_fungal.tsv</strong> for each species of the fungal dataset, at each AuCoMe step (robust orthology, non-robust orthology, and annotation or orthology), several measures were computed, i.e.: the number of reactions, the number of ECs, the number of ECs valided by DeepEC, and ratio number of ECs validated by DeepEC / number of ECs. It was used to design figure S10(a). * <strong>networks_aucome</strong> for each of the 5 fungi (<em>L. bicolor</em>, <em>N. crassa</em>, <em>R. oryzae</em>, <em>S. cerevisiae</em> S288C, and <em>S. pombe</em>), contains a metabolic networks at the PADMet format obtained with AuCoMe. * <strong>networks_gapseq</strong> is composed of 5 subdirectories (one for each fungus). All these subdirectories contain two files about the metabolic networks a obtained with gapseq: – <strong>species-all-Pathways.tbl</strong> encompasses data on pathways at TBL format. – <strong>species-all-Reactions.tbl</strong> includes data on reactions at TBL format. * <strong>networks_modelseed</strong> for each of the 5 fungi (<em>L. bicolor</em>, <em>N. crassa</em>, <em>R. oryzae</em>, <em>S. cerevisiae</em> S288C, and <em>S. pombe</em>), contains two TSV files: – <strong>species.gbk_genome.draftModel-compounds.tsv</strong> includes data on compounds at TSV format. – <strong>species.gbk_genome.draftModel-reactions.tsv</strong> encompasses data on reactions at TSV format. 2/ Content of the <strong>code</strong> repertory<br> It gathers two directories <strong>aucome v0.5.1</strong> and <strong>padmet_v5.0.1</strong>. 2.1/ Content of the <strong>aucome v0.5.1</strong> subdirectory<br> This directory contains a copy of the AuCoMe project on the GitHub site: https://github.com/AuReMe/aucome (downloaded the 15/11/2022). It is composed of two subdirectories and five files:<br> * <strong>LICENCE</strong> licence of the AuCoMe software. * <strong>README.rst</strong> README of the AuCoMe software. * <strong>requirements.txt</strong> contains the list of requires Python packages. * <strong>setup.cfg</strong> contains metadata about AuCoMe package and is used with setup.py to distribute AuCoMe. * <strong>setup.py</strong> contains various information relevant to the AuCoMe package including options and metadata. Then, it is used to distribute AuCoMe with PyPI. It is also used to create an entrypoint when installing it with pip. * <strong>recipes</strong> this subdirectory contains two files:<br> – <strong>Dockerfile</strong> contains instructions to run AuCoMe in a Docker environment.<br> <br> – <strong>Singularity</strong> contains instructions to run AuCoMe in a Singularity container. * <strong>aucome</strong> this directory contains 11 Python files:<br> – <strong>__init__.py</strong> indicates the directory as a python module. – <strong>__main__.py</strong> contains the functions implementing the command-line interface of AuCoMe. – <strong>analysis.py</strong> contains the functions to analyse the AuCoMe results. – <strong>check.py</strong> contains the functions to check the input files. – <strong>compare.py</strong> contains the functions to compare the AuCoMe results between two distinct subgroups. – <strong>orthology.py</strong> contains the functions to propagate reaction through orthology. – <strong>reconstruction.py</strong> contains the functions to perform the reconstruction of draft GSMNs by using Pathway Tools in a parallel implementation.<br> <br> – <strong>spontaneous.py</strong> contains the functions to add spontaneous reactions to some GSMNs if it completes MetaCyc metabolic pathway. – <strong>structural.py</strong> contains the functions to check that no reactions are m



