遇见数据集

Trade-off Predictivity and Explainability for Machine-Learning Powered Predictive Toxicology: An in-Depth Investigation with Tox21 Data Sets

收藏
Figshare2021-01-29 更新2026-04-28 收录
官方服务:

资源简介:

Selecting a model in predictive toxicology often involves a trade-off between prediction performance and explainability: should we sacrifice the model performance to gain explainability or vice versa. Here we present a comprehensive study to assess algorithm and feature influences on model performance in chemical toxicity research. We conducted over 5000 models for a Tox21 bioassay data set of 65 assays and ∼7600 compounds. Seven molecular representations as features and 12 modeling approaches varying in complexity and explainability were employed to systematically investigate the impact of various factors on model performance and explainability. We demonstrated that end points dictated a model’s performance, regardless of the chosen modeling approach including deep learning and chemical features. Overall, more complex models such as (LS-)­SVM and Random Forest performed marginally better than simpler models such as linear regression and KNN in the presented Tox21 data analysis. Since a simpler model with acceptable performance often also is easy to interpret for the Tox21 data set, it clearly was the preferred choice due to its better explainability. Given that each data set had its own error structure both for dependent and independent variables, we strongly recommend that it is important to conduct a systematic study with a broad range of model complexity and feature explainability to identify model balancing its predictivity and explainability.

预测毒理学领域的模型选型,往往需要在预测性能与可解释性之间进行权衡:究竟是牺牲模型性能以换取可解释性,抑或是反之?本研究针对化学毒理学研究中的模型性能与可解释性,系统评估了算法与特征对其产生的影响。我们基于包含65项检测、约7600种化合物的Tox21生物测定数据集,构建了超过5000个模型,并采用7种分子表征作为特征、12种复杂度与可解释性各不相同的建模方法,系统性探究各类因素对模型性能与可解释性的作用效果。研究表明,终点指标决定了模型的性能表现,与所选用的建模方法(包括深度学习方法与化学特征)无关。整体而言,在本次Tox21数据分析中,复杂度更高的模型(如(LS-)支持向量机(SVM)与随机森林(Random Forest))的性能仅略优于线性回归、K近邻(KNN)等更简单的模型。鉴于在Tox21数据集上,性能达标且结构更简单的模型往往更易于解释,因此这类模型凭借更优异的可解释性,显然是更优的选择。由于每个数据集的因变量与自变量均存在独特的误差结构,我们强烈建议:应系统性开展涵盖多种模型复杂度与特征可解释性的研究,以筛选出能够平衡预测性能与可解释性的模型。

创建时间:
2021-01-29
二维码
社区交流群
二维码
科研交流群
商业服务