遇见数据集

Deep Confidence: A Computationally Efficient Framework for Calculating Reliable Prediction Errors for Deep Neural Networks

收藏
Figshare2018-10-30 更新2026-04-29 收录
官方服务:

资源简介:

Deep learning architectures have proved versatile in a number of drug discovery applications, including the modeling of in vitro compound activity. While controlling for prediction confidence is essential to increase the trust, interpretability, and usefulness of virtual screening models in drug discovery, techniques to estimate the reliability of the predictions generated with deep learning networks remain largely underexplored. Here, we present Deep Confidence, a framework to compute valid and efficient confidence intervals for individual predictions using the deep learning technique Snapshot Ensembling and conformal prediction. Specifically, Deep Confidence generates an ensemble of deep neural networks by recording the network parameters throughout the local minima visited during the optimization phase of a single neural network. This approach serves to derive a set of base learners (i.e., snapshots) with comparable predictive power on average that will however generate slightly different predictions for a given instance. The variability across base learners and the validation residuals are in turn harnessed to compute confidence intervals using the conformal prediction framework. Using a set of 24 diverse IC50 data sets from ChEMBL 23, we show that Snapshot Ensembles perform on par with Random Forest (RF) and ensembles of independently trained deep neural networks. In addition, we find that the confidence regions predicted using the Deep Confidence framework span a narrower set of values. Overall, Deep Confidence represents a highly versatile error prediction framework that can be applied to any deep learning-based application at no extra computational cost.

深度学习架构已在包括体外化合物活性建模在内的诸多药物发现应用中展现出优异的通用性。尽管控制预测置信度对于提升药物发现中虚拟筛选模型的可信度、可解释性与实用性至关重要,但针对深度学习网络生成的预测结果估算其可靠性的技术仍未得到充分探索。在此,我们提出Deep Confidence(深度置信框架),该框架可借助深度学习技术Snapshot Ensembling(快照集成)与共形预测(conformal prediction),为单个预测结果计算有效且高效的置信区间。具体而言,Deep Confidence通过记录单个神经网络优化阶段中遍历的各局部极小值处的网络参数,生成深度神经网络集成。该方法可得到一组基础学习器(即快照),这些基础学习器的平均预测能力相当,但针对特定样本会生成略有差异的预测结果。进而,利用基础学习器间的预测差异以及验证集残差,结合共形预测框架计算置信区间。我们使用来自ChEMBL 23的24个多样化IC50数据集开展实验,结果显示快照集成的性能可与随机森林(Random Forest,RF)以及独立训练的深度神经网络集成相媲美。此外,我们发现借助Deep Confidence框架预测得到的置信区域覆盖的数值范围更窄。总体而言,Deep Confidence是一款通用性极强的误差预测框架,可在无需额外计算成本的前提下应用于任何基于深度学习的任务场景。

创建时间:
2018-10-30
二维码
社区交流群
二维码
科研交流群
商业服务