遇见数据集

Supplemental files to the study "Limitations of Current Machine-Learning Models in Predicting Enzymatic Functions for Uncharacterized Proteins"

收藏
Figshare2025-07-18 更新2026-04-28 收录
官方服务:

资源简介:

Thirty to seventy percent of proteins in any given genome have no assigned function and have been labeled as the protein “unknome”. This large knowledge shortfall is one of the final frontiers of biology. Machine-Learning (ML) approaches are enticing, with early successes demonstrating the ability to propagate functional knowledge from experimentally characterized proteins. An open question is the ability of machine-learning approaches to predict enzymatic functions unseen in the training sets. Using a set of E. coli unknowns, we evaluated the current state-of-the-art machine-learning approaches and found that these methods currently lack the ability to integrate scientific reasoning into their prediction algorithms. While human annotators are able to leverage the plethora of genomic data in making plausible predictions into the unknown, current ML methods not only fail to make novel predictions but also make basic logic errors in their predictions. This underscores the need to further develop ML methods and to deploy deterministic approaches to test for ‘hallucinations’ and other aberrant behavior in the current generation of predictive modeling. eXplainable AI (XAI) analysis revealed that noisy, ambiguous, or low contribution profiles across the protein sequence are strong indicators of unreliable predictions, enabling systematic identification of potential errors and elimination of most uncertain predictions. These findings demonstrate the value of integrating XAI-driven filtering strategies to improve the reliability of machine-learning-based protein annotation.

任何给定基因组中30%至70%的蛋白质尚未被赋予明确功能,被归类为蛋白质‘未知组(unknome)’。这一巨大的知识缺口是生物学领域的最后前沿阵地之一。机器学习(Machine-Learning, ML)方法颇具吸引力,早期研究已证明其可从经实验表征的蛋白质中传递功能认知。一个悬而未决的问题是,机器学习方法能否预测训练集中未出现过的酶功能。我们利用大肠杆菌(E. coli)的未知蛋白质集,对当前最先进的机器学习方法进行了评估,发现这些方法目前尚无法将科学推理融入预测算法。尽管人类注释者能够借助海量基因组数据对未知蛋白质做出合理预测,但当前的机器学习方法不仅无法做出新颖预测,还会在预测中出现基本逻辑错误。这凸显了进一步优化机器学习方法,并部署确定性方法以检测当前预测建模生成‘幻觉’及其他异常行为的必要性。可解释人工智能(eXplainable AI, XAI)分析显示,蛋白质序列上存在噪声、歧义或低贡献度的特征分布,是预测不可靠的强烈信号,这使得我们能够系统性识别潜在错误并剔除大部分不确定性预测。这些研究结果证明,整合可解释人工智能驱动的过滤策略,可提升基于机器学习的蛋白质注释的可靠性。

创建时间:
2025-07-18
二维码
社区交流群
二维码
科研交流群
商业服务