Protein structure network (PCN) and residue interaction graph (RIG) data for Anti-CRISPR protein analysis
收藏资源简介:
This contains majority of the data files used in the paper - "Revisiting graph-based approaches for small protein analysis: Insights from anti-CRISPR protein networks". Contains protein contact networks (PCNs) and residue interaction graphs (RIGs). The Github associated with this project is AcrNetworks-clean. This work uses data from the Protein Data Bank. Data provided by the RCSB PDB (rcsb.org) [1]. The Citation: H.M. Berman, J. Westbrook, Z. Feng, G. Gilliland, T.N. Bhat, H. Weissig, I.N. Shindyalov, P.E. Bourne. (2000) The Protein Data Bank Nucleic Acids Research, 28: 235-242. It also utilizes Arpeggio to gather infromation about the interatomic contacts to help with the production of the RIG graphs. Citation: Harry C Jubb, Alicia P Higueruelo, Bernardo Ochoa-Montaño, Will R Pitt, David B Ascher, Tom L Blundell, Arpeggio: A Web Server for Calculating and Visualising Interatomic Interactions in Protein Structures. Journal of Molecular Biology, Volume 429, Issue 3, 2017, Pages 365-371, ISSN 0022-2836. Original Folder Layout within Code - Not all files are included due to sizing constraints, see "Not included here but available upon request". Please contact me directly for the files. Important Files - 3 main Acr data - centrality and node information for both PCNs and RIGs of AcrIF1, AcrIIA1, and AcrVIA1 3 main Acr PDB files - PDB files associated with AcrIF1, AcrIIA1, and AcrVIA1 Analysis Data - Organism distribution and amino acid analysis distribution Cleaned Centralities - Centrality information for all RIGs for the PISCES data set Cleaned JSONs - JSON representation for all RIGs for the PISCES data set. Generated by Arpeggio. Cleaned PDBs - the converted mmCIF files from the PDBe-Arpeggio cleaning process. Cleaned RIGs - RIGs created using cleaned PDBs Cleaned SASAs - SASA and RSA information for each protein in PISCES data set mmCIF files - PDBe-Arpeggio cleaning returns a mmCIF, which is then converted to a PDB. PISCES - This folder contains a mix of used and unused files of the paper and remnants of troubleshooting and testing. Some folders that are explicitly used were copied into Important Files for keeping it clear what was and was not used within the paper. Many of the .py files in the code path to this folder. [deprecated] implies the file was a testing file or from an old or failed run. Centralities - same as Cleaned Centralities (Important Files) CIFs - same as mmCIF files (Important Files) GraphMLs - same as Cleaned RIGs (Important Files) known acr data - PCNs and RIGs for AcrIF1, AcrIIA1, and AcrVIA1 PDBs_compressed - compressed versions of the PDBs (raw) SASAs - same as Cleaned SASAs PISCES_culled.fasta - the PISCES call returned some PDBs that had to be filtered for size or for lacking data, these are the remaining PDBs kept for use PISCES_data_info_mutant.csv - contains information about proteins which are considered 'mutants' out of the used PDBs PISCES_data_info - contains information about all proteins pulled (mutants and non-mutants) PISCES_pdbs.txt - a text file with all of the PDB ids used Not included here but available on request: Cleaned JSONs - JSON representation for all RIGs for the PISCES data set. Generated by Arpeggio. (Roughly 33 GB generated size. - Not provided but generated using the "parallel_arpeggio.py" file.) CIFs_large: unused data due to size constraints JSONs: [deprecated] same as Cleaned JSONs (Important Files) PDBs: PDB files pulled using Protein Data Bank script (batch_download.sh) PDBs_cleaned: [deprecated] from an old run



