The Experimental Uncertainty of Heterogeneous Public Ki Data
收藏资源简介:
The maximum achievable accuracy of in silico models depends on the quality of the experimental data. Consequently, experimental uncertainty defines a natural upper limit to the predictive performance possible. Models that yield errors smaller than the experimental uncertainty are necessarily overtrained. A reliable estimate of the experimental uncertainty is therefore of high importance to all originators and users of in silico models. The data deposited in ChEMBL was analyzed for reproducibility, i.e., the experimental uncertainty of independent measurements. Careful filtering of the data was required because ChEMBL contains unit-transcription errors, undifferentiated stereoisomers, and repeated citations of single measurements (90% of all pairs). The experimental uncertainty is estimated to yield a mean error of 0.44 pKi units, a standard deviation of 0.54 pKi units, and a median error of 0.34 pKi units. The maximum possible squared Pearson correlation coefficient (R2) on large data sets is estimated to be 0.81.
计算机模拟(in silico)模型所能达到的最高精度,取决于实验数据的质量。据此,实验不确定性为模型可实现的预测性能划定了天然上限。若模型的预测误差小于实验不确定性,则必然属于过度训练的情况。因此,对实验不确定性进行可靠的估计,对于所有计算机模拟模型的开发者与使用者而言都具有极高的重要性。本研究针对存入ChEMBL数据库的数据开展可重复性分析,即评估独立测量值的实验不确定性。由于ChEMBL数据库中存在单位转录错误、未区分的立体异构体,以及单条测量结果的重复引用(占所有数据对的90%),因此需要对数据进行严格的筛选处理。经估计,本次实验不确定性对应的平均误差为0.44 pKi单位,标准差为0.54 pKi单位,中位数误差为0.34 pKi单位。经估计,大型数据集上可达到的皮尔逊相关系数平方(R²)的最大值为0.81。




