Data for "Training data composition affects performance of protein structure analysis algorithms" by A. Derry, K. A. Carpenter, & R. B. Altman
收藏资源简介:
<strong>Description</strong> This repository contains all data used in "Training data composition affects performance of protein structure analysis algorithms", published in the Pacific Symposium on Biocomputing 2022 by A. Derry, K. A. Carpenter, & R. B. Altman. The data consists of the following files: ema_zenodo_data.tar.gz: train, validation, and test splits for Estimation of Model Accuracy task, in LMDB format design_zenodo_data.tar.gz: train, validation, and test splits for Protein Sequence Design task, in JSON format enz_cat_res_zenodo_data.tar.gz: train, validation, and test splits for Catalytic Residue and Enzyme Prediction task, in TF record format Details on dataset construction can be found in our paper and dataloaders can be found in our Github repo. <strong>Reference</strong> A. Derry*, K. A. Carpenter*, & R. B. Altman, "Training data composition affects performance of protein structure analysis algorithms", 2021. <strong>Dataset References</strong> Datasets used were derived from the following works: Kryshtafovych, A., Schwede, T., Topf, M., Fidelis, K., & Moult, J. (2019). Critical assessment of methods of protein structure prediction (CASP)—Round XIII. In <em>Proteins: Structure, Function and Bioinformatics</em> (Vol. 87, Issue 12, pp. 1011–1020). https://doi.org/10.1002/prot.25823 Ingraham, J., Garg, V. K., Barzilay, R., & Jaakkola, T. (2019). <em>Generative Models for Graph-Based Protein Design</em>. https://openreview.net/pdf?id=SJgxrLLKOE Furnham, N., Holliday, G. L., de Beer, T. A. P., Jacobsen, J. O. B., Pearson, W. R., & Thornton, J. M. (2014). The Catalytic Site Atlas 2.0: cataloging catalytic sites and residues identified in enzymes. <em>Nucleic Acids Research</em>, <em>42 </em>(Database issue), D485–D489.



