Supplementary Data for "Exploration of the structural and functional diversity in the metamorphic RfaH subfamily"
收藏资源简介:
This Supplementary Data is provided as a compressed ZIP file, which contains the following folders: RfaH_FL: This folder contains AlphaFold2 (AF2) protein structure predictions of the full-length protein sequences of RfaH orthologs using LocalColabFold, a local version of ColabFold v1.5.5, with default parameters (1 seed, 5 models, alphafold2_ptm model for monomers, 3 recycles, no templates, multiple sequence alignments [MSA] generated using MMseqs2). The folder contains several files and subfolders: PDB files: Best ranked predicted structure of E. coli αRfaH (arfah.pdb) and its α-folded C-terminal domain (αCTD, actd.pdb) using ColabFold, and the experimental structure of the β-folded CTD of E. coli (βCTD, PDB ID 2LCL, bctd.pdb), which were used for computing the structural similarity of the CTD between the predicted structures and the αCTD and βCTD of E. coli. interpro: Subfolder containing all the structures predicted for RfaH orthologs from InterPro, family entry IPR010215, 3,058 sequences, accessed on January 3 2024. colabfolddb: Subfolder containing all the structures predicted for RfaH orthologs from ColabFoldDB, 468 sequences, accessed on January 16 2026 via MMseqs2 using the sequence of E. coli RfaH (UniProt accession ID P0AFW0) as input, and filtered based on a minimum sequence coverage of 95%, a bit-score of 140, a maximum e-value of 1e-30, and elimination of redundant sequences from InterPro database. cluster: Subfolder containing all the structures predicted for RfaH orthologs from a reported list of 416 RfaH sequences (Cluster 1) curated by Hidden Markov Models, with known genomic context, that were classified by Markov clustering and include E. coli RfaH. MSA files: MSAs in FASTA format, generated using MAFFT and including the sequence of E. coli RfaH, of the protein sequences from the InterPro (mafft_EcRfaH_all_sequences_interpro.fasta), ColabFoldDB (mafft_EcRfaH_all_sequences_colabfolddb.fasta) and Cluster (mafft_EcRfaH_all_sequences_cluster.fasta) databases. RfaH_CTD: This folder contains AF2 protein structure predictions of the CTD protein sequences of RfaH orthologs using LocalColabFold with default parameters. The folder contains the same structure of files (PDBs and MSAs) and subfolders (interpro, colabfolddb, cluster) than RfaH_FL. TMD_FL_decoys: This folder contains a set of structural decoys generated via implicit-solvent Targeted Molecular Dynamics (TMD) of the fold-switch of the CTD of E. coli RfaH in the context of the full-length protein using AMBER22. Simulations systems were prepared based on structures of full-length E. coli RfaH in the autoinhibited and active states using PDB IDs 5OND (NTD), 5OUG (αCTD) and 2LCL (βCTD), with missing atoms being added using MODELLER v9.23 and minimizing the structures in explicit solvent (TIP3P water molecules, ff99SB-ILDN force field) using GROMACS v4.5.3 using the steepest descent method. The simulation systems were prepared using AmberTools v22.3 with the ff19SB force field and the GB-Neck 2 implict solvent model with mbondi3 intrinsic radii. Each system was minimized for 2,500 steps using a steepest descent gradient, followed by 47,000 steps of conjugate gradient, before performing 300 short (100 ps) TMD simulations from αRfaH to βRfaH and 300 TMD simulations in the opposite direction at 298 K. The forces for the TMD potential were randomly selected between 1.0×10⁻⁴ to 3×10⁻² kcal/(mol·Å²) applied over the backbone atoms of residues 110-162, with a target RMSD of 0 Å to the target structure. Principal Component Analysis was performed to choose ~100 TMD trajectories in each direction that represented the full fold-switch landscape, leading to a curated dataset of 100 TMD trajectories for the simulations from αRfaH to βRfaH (3,750 decoy structures) and 99 TMD trajectories for the simulations from βRfaH to αRfaH (3,700 decoy structures). Then, the TMD decoys were analyzed using AF2Rank with default parameters (1 recycle, alphafold_ptm model and 1 output model per input decoy), using both the AMBER-minimized αRfaH and βRfaH states of E. coli RfaH as reference native states. The folder contains several files and subfolders: input_PDBs: Initial structures generated using experimental structures of E. coli RfaH and MODELLER and minimized in GROMACS. PCA_aRfaH: 3,750 decoy structures from 100 TMD simulations from αRfaH to βRfaH, with 1 trajectory being represented every 50 frames. PCA_bRfaH: 3,700 decoy structures from 99 TMD simulations from βRfaH to αRfaH, with 1 trajectory being represented every 50 frames. PDB_files: AMBER-minimzed structures of E. coli αRfaH (aRfaH.pdb) and βRfaH (bRfaH.pdb) to be used as native states for AF2Rank. AF2Rank_A-B_aRfaH: 7,500 AF2Rank output structures after analyzing the decoys of the αRfaH to βRfaH TMD trajectories using E. coli αRfaH as native state. AF2Rank_A-B_bRfaH: 7,500 AF2Rank output structures after analyzing the decoys of the αRfaH to βRfaH TMD trajectories using E. coli βRfaH as native state. AF2Rank_B-A_aRfaH: 7,000 AF2Rank output structures after analyzing the decoys of the βRfaH to αRfaH TMD trajectories using E. coli αRfaH as native state. AF2Rank_A-B_aRfaH: 7,000 AF2Rank output structures after analyzing the decoys of the βRfaH to αRfaH TMD trajectories using E. coli βRfaH as native state. notebooks: This folder contains a collection of 4 Jupyter Notebooks for execution on Google Colab to reproduce the main findings of our work. The notebooks are meant to be run in sequence and correspond to: 01_Analysis_AF2_RfaH_CTD: Jupyter Notebook for the structural analysis of the AF2 predicted structures of the CTD of RfaH orthologs from InterPro, ColabFoldDB and Cluster, using TMalign to calculate TM-score and RMSD of the CTD of the predicted structures against the reference αCTD and βCTD structures of E. coli RfaH, and STRIDE v1.6.4 to calculate the per-residue secondary structure content, as well as extracting the CTD predicted Local Distance Diference Test (pLDDT), a confidence metric of the quality of the predicted models. This data is compiled into a CSV file for structure categorization into metamorphic αRfaH (percentage of α-helices > 32.5% and < 2.5% of β-strands), monomorphic βRfaH (percentage of β-strands > 30.0% and < 2.5% of α-helices) or mixed secondary structure (percentage of α-helices and β-strands > 2.5%). It also generates images for visualization of the results. 02_Analysis_AF2_RfaH_FL: Jupyter Notebook for the structural analysis of the AF2 predicted structures of the full-length sequences of RfaH orthologs from InterPro, ColabFoldDB and Cluster. It is similar to the previous notebook, but it also calculates the sequence identity of the best predicted structures in each category, the Cβ-Cβ distance between residues in equivalent positions to the E48-R138 interdomain salt bridge in E. coli RfaH, based on the MAFFT alignments of all RfaH orthologs for the three databases, and uses Logomaker to calculate the sequence logo of these residues. 03_AF2Rank_Analysis_TMD: Modified version of the AF2Rank Jupyter Notebook to analyze multiple structures (here, the TMD-generated decoys) against two native states and save the output AF2Rank parameters and structures into Google Drive. WARNING: Running this notebook will overwrite the existing files from folders AF2Rank_A-B_aRfaH, AF2Rank_A-B_bRfaH, AF2Rank_B-A_aRfaH and AF2Rank_A-B_aRfaH. 04_AF2Rank_Plot_Results: Jupyter Notebook to analyse the results from AF2Rank. It calculates the backbone RMSD of the input TMD decoys and the output AF2Rank structures using cpptraj from AmberTools, performs a clustering analysis of the AF2Rank outputs by k-means and exports the representative medoid structures from each cluster (here, using k = 5) to a separate folder in Google Drive. It also generates several plots of the TMD trajectories, the mapping between the input TMD decoys and the output AF2Rank structures and directional path plots to check for hysteresis. analysis: This folder contains the results of the analysis of the AF2 predictions of RfaH orthologs and the TMD trajectories and AF2Rank output structures using the Jupyter Notebooks mentioned above. It contains several subfolders: CF_RfaH_FL: Analysis of the AF2 predicted structures of full-length RfaH orthologs from InterPro, ColabFoldDB and Cluster. It contains 4 CSV files: rfaH_raw_structural_data: CSV file with the full-length and CTD RMSD and TM-score of all predictions using TMalign, and the percentage of α-helical and β-strand secondary structure using STRIDE. pdb_categorization_all: CSV file with the categorization of each structure as metamorphic, monomorphic, mixed or none, as well as calculating the sequence identity percentage based on the MAFFT alignments and the pLDDT for the CTD of each structure. category_pctid_unique: CSV file with the best predicted structure per category for each RfaH ortholog. If the 5 structures predicted for each RfaH ortholog fall into multiple categories, the best predicted structure per category is added to the list. This file is used to calculate the sequence identity percentage and pLDDT per classification category. rfah_sidechain_distances_master.csv: It contains the Cβ-Cβ distance between residues in equivalent positions to the E48-R138 interdomain salt bridge in E. coli RfaH, based on the MAFFT alignment, as well as the residues in each position. CF_RfaH_CTD: Analysis of the AF2 predicted structures of the isolated CTD of the RfaH orthologs from InterPro, ColabFoldDB and Cluster. It contains 3 CSV files, namely rfaH_raw_structural_data, pdb_categorization_all and category_pctid_unique, similar to the previous folder. TMD_AF2Rank: Analysis of the input TMD decoys in each direction (Results_TMD_A-B.csv and Results_TMD_B-A.csv) and output AF2Rank structures in terms of confidence metrics (predicted alignment error [PAE], predicted TM-score [pTM] and pLDDT), AF2Rank-specific confidence metrics (input TM-score of the decoy structure [tm_i], output TM-score between the decoy structure versus the reference native structure [tm_o], TM-score between the input structure and the AF2 output structure [tm_io], composite score [pLDDT×pTM×tm_io]), as well as the Cα RMSD for the CTD of the predicted structures from AF2Rank (RMSD_AF2R) and the TMD input decoys (RMSD) vs the reference native states. It contains a subfolder: Representative_Medoids: Folder with AF2Rank output structures of the representative medoids after k-means clustering (k = 5). BRIEF INSTRUCTIONS To reproduce the data analysis presented in the manuscript associated to this work, you must 1) Download the ZIP file and uncompress it 2) Save the uncompressed folder into the main folder of your Google Drive with the name "AF2_RfaH_2026) 3) Otherwise, you might have to change several path variables in the Jupyter Notebooks for execution in Google Colab.



