Data for RAPPPID: Towards Generalisable Protein Interaction Prediction with AWD-LSTM Twin Networks
收藏资源简介:
Data for RAPPPID, a method for the Regularised Automative Prediction of Protein-Protein Interactions using Deep Learning. These datasets are in a format that RAPPPID is ready to read.<br> <br> <strong>Comparatives Dataset</strong><br> These datasets were derived from the STRING v11 <em>H. sapiens</em> dataset, according to the C1, C2, and C3 procedures outlined by Park and Marcotte, 2012. Negative samples are sampled randomly from the space of proteins not known to interact. See Szymborski & Emad for details.<br> <br> <strong>Repeatability Datasets</strong><br> The following datasets are all derived from STRING in the manner as the comparatives dataset, but three different random seeds are used for drawing proteins.<br> <br> <strong>References</strong><br> Park,Y. and Marcotte,E.M. (2012) Flaws in evaluation schemes for pair-input computational predictions. Nat Methods, 9, 1134–1136. Szklarczyk, D., Gable, A. L., Lyon, D., Junge, A., Wyder, S., Huerta-Cepas, J., Simonovic, M., Doncheva, N. T., Morris, J. H., Bork, P., Jensen, L. J., and Mering, C. (2019). String v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets. Nucleic Acids Research, 47(D1), D607–D613.<br> <br> Szymborski,J. and Emad,A. (2021) RAPPPID: Towards Generalisable Protein Interaction Prediction with AWD-LSTM Twin Networks. bioRxiv https://doi.org/10.1101/2021.08.13.456309



