Sensitivity Datasets - Leveraging Implicit Knowledge In Neural Networks For Functional Dissection And Engineering Of Proteins
收藏资源简介:
<strong>Leveraging Implicit Knowledge in Neural Networks for Functional Dissection and Engineering of Proteins</strong> The Sensitivity datasets cover more than 2000 proteins and are structured as follows. It is uploaded as tar.gz. and structured in three separate directories. mean_pdb/ contains the proteins for the proteins used to analyze the sphere variances, correlation with information content and correlations between GO terms (Figure 2c-e) and the ligand binding (Figure 3 a-c, Supplementary Figure 3). mean_examples/ contains the proteins used for inferring the protein-receptor hybrids by the Hahn lab (Figure 5) with_biological_activity/ contains the ERK2 data (Figure 3), the spCas9 data (Supplementary Figure 5) and the AcrIIA4 data (Figure 6)<br> binding_activities.csv contains the pdb identifiers and ligand descriptions for Figure 3, Supplementary Figure 3 and is needed for distance_to_ligand.py <strong>Important</strong> The biological activity data for spCas9 is from Brenan et al.<sup>1</sup>: Supplementary Table 1. We used column ‘dox_average’, here 'mean_dox_average'. The biological activity data for ERK2 is from Oakes et al.<sup>2</sup>: Supplementary Table 2. We used column ‘fold_change’ log2-transfomred, here 'mean_log2_fold_change'. The sequences and secondary structure information were downloaded from the RCSB Protein Databank and are available here: https://cdn.rcsb.org/etl/kabschSander/ss_dis.txt.gz This URL can be found with some explanation at http://www.rcsb.org/pdb/static.do?p=download/http/index.html The secondary structure annotation relies on the DSSP Algorithm by Kabsch and Sander<sup>3</sup> <strong>The files are tab-separated and contain the following columns:</strong> <strong>Pos</strong> Position in the sequence, starting from zero <strong>AA</strong> Amino acid in that position <strong>sec</strong> Secondary structure as annotated in the RCSB Protein Databank <strong>dis</strong> if a region has not been experimentally observed (sometimes explains mismatches with crystal structures) <strong>GO:_______</strong> Sensitivity for that GO term <strong>svar_GO:_______</strong> Shere Variance of the sensitivity for that GO term <strong>ic</strong> Information content, based on Pfam seed alignment <strong>svar_n_neighbours</strong> number of residues in the sphere used to calculate the sphere variance <strong>svar_d_center</strong> Distance to the center of mass of the chain that was analyzed <strong>Others</strong> refer to biological activity data, depend on the source <strong>References</strong> Brenan, L. et al. Phenotypic Characterization of a Comprehensive Set of MAPK1/ERK2 Missense Mutants. Cell Rep 17, 1171-1183, doi:10.1016/j.celrep.2016.09.061 (2016). Oakes, B. L. et al. Profiling of engineering hotspots identifies an allosteric CRISPR-Cas9 switch. Nat Biotechnol 34, 646-651, doi:10.1038/nbt.3528 (2016). Kabsch, W. & Sander, C. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features. Biopolymers 22, 2577-2637, doi:10.1002/bip.360221211 (1983).



