UdonPred Data
收藏资源简介:
Datasets used to train the UdonPred Method.UdonPred is a predictive model for protein disorder, leveraging embeddings from the pre-trained protein language model ProstT5. This model variant has been trained using a novel dataset of nuclear magnetic resonance (NMR) spectroscopy chemical shifts from the Biological Magnetic Resonance Data Bank (BMRB). It employs the TriZOD scoring scheme, which quantifies disorder by assigning G-scores that identify deviations from the chemical shifts expected in random coil configurations. UdonPred was subsequently fine-tuned using data from the DisProt database.UdonPred is available at: https://github.com/JSchlensok/udonpred<br>TriZOD G-Scores method is available at: https://github.com/MarkusHaak/trizod
用于训练UdonPred方法的数据集。UdonPred是一款蛋白质无序性预测模型,利用预训练蛋白质语言模型ProstT5生成的嵌入(embedding)特征开展工作。该模型变体基于生物核磁共振数据银行(Biological Magnetic Resonance Data Bank, BMRB)收录的新型核磁共振(NMR)光谱化学位移数据集完成训练。其采用TriZOD评分体系,通过赋予G值(G-scores)量化蛋白质无序性,以此识别与随机卷曲构象下预期化学位移的偏差。后续研究人员借助DisProt数据库的数据对UdonPred进行了微调。UdonPred的开源代码仓库地址为:https://github.com/JSchlensok/udonpred TriZOD G值评分方法的开源代码仓库地址为:https://github.com/MarkusHaak/trizod




