Sars-CoV-2 structures -- sequence-to-alignments derived from PDB and from PSSH2 plus dark regions
收藏资源简介:
<strong>Protein sequence and structure data</strong> This data set contains data from Uniprot (in the files called protein_sequence, protein_synonyms, protein_namess, organism_synonysm) and PDB (in the files called PDB and PDB_chain) as used by the Aquaria web resource at the time of download (2021-06-28). <strong>The PSSH2 data set</strong><br> <br> PSSH2 is a database of protein sequence-to-structure homologies based on HHblits, an alignment method employing iterative comparisons of hidden Markov models (HMMs). To ensure the highest possible final alignment quality for matches in Aquaria using HHblits, we first calculate HMM profiles for each unique PDB sequence (PDB_full) and also for each unique Swiss-Prot sequence. We generated PSSH2 using HHblits to find similarities between HMMs from PDB and HMMs from UniProt sequences. <strong>Calculating PSSH2</strong> The main bunch of Swissprot and PDB data was downloaded in October 2020, but incremental updates, especially as related to Covid19 were added until April 2021.<br> Generating PSSH2: We used Uniclust30 from HH-suite, a database of non-redundant UniProt sequence clusters in which the highest pairwise sequence identity between clusters was 30% (http://gwdu111.gwdg.de/~compbiol/uniclust/2020_03/UniRef30_2020_03_hhsuite.tar.gz). The HHblits code and the code for running the calculations was retrieved from git (https://github.com/soedinglab/hh-suite.git and https://github.com/aschafu/PSSH2.git respectively) at the respective time of calculation in the timeframe until April 2021. <br> <strong>PDB based sequence-to-structure alignments</strong> In addition to the PSSH2 data, new PDB structures were retrieved based on the primary accession of the proteins, by querying for all chains in all PDB entries with exact matches using the sequence cross references records given in PDB. Sequence-to-structure alignments were then created, again based on information provided in each PDB entry. These are contained in the PDBchain data. This data covers sequences and PDB structures in the timeframe until June 2021.



