File S1 - Prediction and Analysis of Canonical EF Hand Loop and Qualitative Estimation of Ca<sup>2+</sup> Binding Affinity
收藏资源简介:
File S1 includes the following: Figure S1. a) Plot of affinity vs. PSSM for the test data set (D5). The calculated correlation coefficient obtained was 0.61 using [41] amino acid frequencies. Figure S2. The isothermal titration calorimetric analysis of Ca2+-binding to apo-EhCaBPs.ITC experiments were carried out as described under “Materials and Methods”. Plot of heat absorbed/released (In kcal mol−1) per injection of CaCl2 as a function of molar ratio of Ca2+: protein at 25°C is shown. For all titrations, the top panels represent the raw data (power: time) and the bottom panels represent integrated binding isotherms. The solid line represents the best nonlinear fit to the experimental data. Binding isotherm for A: EhCaBP3; B: EhCaBP4; C: EhCaBP5; D: EhCaBP6 and E: EhCaBP7. Thermodynamic parameters obtained are summarized in Table 1. Figure S3. ROC plots of AC&CC, AC&HC, AC&HC&HYC, AC&HYC&CC and AC&HYC for the datasets D5–D7 set. Receiver operating characteristic (ROC) plot used for depicting relative trade-offs between true positive and false positives. The corresponding AUC value of each model is shown in brackets. Figure S4. Schematic representation of the procedure for model development and feature selection for EF-hand loop region prediction and estimation of binding affinity and its web implementation. The procedure is explained in detail in the “Methods” section. A). A group of sequences with known EF-hand structural motifs were downloaded and further classified into two groups after removing redundant sequences using CD-HIT. The sequences were further converted into binary and amino acid composition (AAC) profiles for SVM input. Models were generated using LIBSVM and were tested on all the datasets (D3–D6) and further validated by scanning the E. histolytica proteome. B). Non-redundant sequences of EF-hand loops from known structures were classified into two groups on the basis of scores obtained from position-specific scoring metrics. The sequences were then converted into binary, AAC and different amino acid indices patterns. We have generated both standalone and combinations of features (2, 3, 4, 5) using a Perl script written in-house. The input vectors were trained using LIBSVM and cudized LIBSVM and selected on the basis of their performance on experimental datasets using 5-fold cross validation accuracy threshold >70 %. The best performing models selected from screening were further validated using three different experimentally derived datasets on EF hand motifs. The final step involved web implementation of the best (AC&HC) model. Table S1. The χ2 value for each amino acid residue is estimated with one degree of freedom and significance level P = 0.001. The Σχ2 values are estimated with 19 degrees of freedom and significance level P<0.001. The expected (Exp) and observed (Obs) values and the corresponding χ2 values for amino acid residues and the Σχ2 values for those positions that do not reach 10.8 and 43.8 (for one and 19 degrees of freedom, respectively) are given more significance. Table S2. Test Dataset: Summary of EF hand loops obtained from the literature and their macroscopic binding constant along with CAL-EF-AFi predictions (D5). The classification details with supportive binding constants are listed under “Author's Note”. (Red-colored affinities are the false negative affinity predictions, and turquoise-colored sequences are the false negative EF loop predictions). Table S3. Independent dataset (D6) summary of EF hand loops obtained from Boguta, et al., 1988 [59]. The table contains average binding constants of Ca2+ for troponin C superfamily (TnC) proteins from experimental data reported by various laboratories. The classification details with supportive binding constants are listed under “Author's Note”. (Red-colored affinities are the false positive predictions). Table S4. Validation dataset summary of EF-hand loops obtained from ITC studies of CaBPs from E. histolytica and their macroscopic binding constant according to CAL-EF-AFi's predictions (D7). The classification details with supportive binding constants are listed under “Author's Note” (Red-colored affinities are the false positive predictions). Table S5. Predictions of putative EF hand-containing calcium-binding protein and their calcium-binding affinities from the E. histolytica proteome. Table S6. The performance and comparison of CAL-EF-AFi with PFAM and Calpred on the E. histolytica proteome. Listed are the sequences predicted by CAL-EF-AFi followed by PFAM-based HMM model prediction and CalPred's predictions. (Legends for CAL-EF-AFi's prediction: number of Ca2+-binding loop sequence prediction, residue number followed by sequence and SVM scores; Legends for PFAM predictions: red-colored region is the loop region predicted, followed by the E-value for the sequence; Legends for CalPred predictions: X: Non-Binding region C: Calcium Binding region). Table S7. Calcium-binding EF-hand protein sequences in FASTA format at 60% sequence redundancy with EF-hand loop region residues labeled in lower case letters. (D1). Table S8. The list of 12-mer sequences from non-binding regions of calcium-binding EF-hand proteins greater than 60% sequence redundancy. Table S9. The training data used for estimation of binding affinity were taken from the RCSB based on PSSM scores obtained from the EF-hand loop region. The positive dataset (D3) consisted of one hundred forty four 12-mer sequences and there were 124 sequences in the negative dataset (D4). Table S10. The redundant set of PDB ids of EF hand-containing calcium-binding proteins. The sequences taken from the RCSB were further processed using CD-HIT and the list if the sequences with different threshold are listed in Table S11. Table S11. The sequence-wise classification of data obtained from PROSITE and RCSB- The data was further processed by using CD-HIT at 90%, 70%, 60%, 50% sequence redundancy cutofffor classification of EF-hand loop Ca2+-binding and non-binding region. (DOC)



