遇见数据集

Mutational data for protein solubility

收藏
知名数据库2026-06-02 收录
官方服务:

资源简介:

SoluProtMutDB is a comprehensive, manually curated database of the protein solubility data. Low protein solubility presents major challenges to industrial applications and is often reported to be behind many human diseases. Understanding how mutations affect protein solubility can therefore help elucidate the mechanisms associated with the development of human diseases and better utilize protein engineering in industrial applications. Multiple factors may play a role here: the presence of a chaperone or co-factor required for correct protein folding, unnatural physiological conditions such as high temperature, pH, or protein concentration, tendency of a protein to aggregate due to aggregation-prone regions, etc. The predictive power of the existing protein engineering tools is often compromised by limited experimental data available for rigorous training and testing of solubility predictions. The published data used for solubility prediction upon mutation are usually scattered in the literature and had to be collected manually. The goal of the SoluProtMut database is to collect the reported evidence of solubility changes upon mutations from published sources to guide future protein engineering effort in producing soluble protein variants. The database currently contains data previously used for training Machine Learning-based predictors, such as PON-Sol, CamSol, AGGRESCAN3D, OptSolMut, as well as recently published datasets. We are providing manually curated and reliable data in the standardized format which are pre-processed for machine learning applications.

SoluProtMutDB是一个经人工审编(manually curated)的综合性蛋白质溶解度数据库。蛋白质溶解度偏低是工业应用面临的重大挑战,同时据报道也是诸多人类疾病的致病诱因之一。因此,明晰突变对蛋白质溶解度的影响机制,不仅有助于解析人类疾病发生的相关分子机制,还能推动蛋白质工程技术在工业场景中的优化应用。该过程受多种因素调控:包括蛋白质正确折叠所需的分子伴侣(chaperone)或辅因子(co-factor)的存在、高温、异常pH值或蛋白质浓度等非生理条件,以及蛋白质因聚集倾向区域(aggregation-prone regions)引发的聚集趋势等。现有蛋白质工程工具的预测能力往往受限于可用实验数据不足,难以开展严格的溶解度预测模型训练与测试。此前用于突变后溶解度预测的公开数据通常分散于各类学术文献中,需通过人工方式逐一收集整理。SoluProtMutDB数据库的构建目标,是从已发表的文献资料中搜集突变引发蛋白质溶解度变化的相关实证数据,为未来通过蛋白质工程制备可溶性蛋白变体的研究提供指导。当前该数据库收录了此前用于训练基于机器学习(Machine Learning)的预测模型的数据集,如PON-Sol、CamSol、AGGRESCAN3D、OptSolMut,以及近年新发表的相关数据集。我们以标准化格式提供经人工审编的可靠数据,这些数据已完成预处理,可直接用于机器学习相关应用。

提供机构:
马萨里克大学
搜集汇总
数据集介绍
Mutational data for protein solubility 数据集图片
背景与挑战
背景概述
SoluProtMutDB是一个手动整理、全面的数据库,专门收集突变导致蛋白质溶解度变化的报告证据,以指导蛋白质工程中可溶性变体的生产。该数据库包含用于机器学习预测工具训练的数据,并以标准化格式提供,适用于机器学习应用。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务