Trade-off Predictivity and Explainability for Machine-Learning Powered Predictive Toxicology: An in-Depth Investigation with Tox21 Data Sets.

Trade-off Predictivity and Explainability for Machine-Learning Powered Predictive Toxicology: An in-Depth Investigation with Tox21 Data Sets.
复制标题

DOI:
10.1021/acs.chemrestox.0c00373
复制
发表时间:
2021-02-15
影响因子:
4.1
通讯作者:
Tong W
Tong W
中科院分区:
医学3区
文献类型:
--
作者:
Wu L;Huang R;Tetko IV;Xia Z;Xu J;Tong W

文献摘要

参考文献

被引文献

相似文献

在预测毒理学中选择模型通常涉及预测性能和可解释性之间的权衡:我们应该牺牲模型性能来获得可解释性,还是反之亦然?在这里,我们提出了一个全面的研究,以评估算法和功能对模型性能的影响,在化学毒性研究。我们对Tox 21生物测定数据集进行了超过5000个模型,其中包括65种测定和~7600种化合物。七个分子表示的功能和12个建模方法不同的复杂性和可解释性,系统地研究了各种因素对模型性能和可解释性的影响。我们证明了端点决定了模型的性能,无论选择何种建模方法,包括深度学习和化学特征。总体而言,在所呈现的Tox 21数据分析中,更复杂的模型(如(LS-)SVM和随机森林)的表现略好于更简单的模型(如线性回归和KNN)。由于具有可接受性能的更简单模型通常也易于解释Tox 21数据集,因此由于其更好的可解释性,它显然是首选。鉴于每个数据集对于因变量和自变量都有自己的误差结构,我们强烈建议对广泛的模型复杂性和特征可解释性进行系统研究,以确定平衡其预测性和可解释性的模型。
Selecting a model in predictive toxicology often involves a trade-off between prediction performance and explainability: should we sacrifice the model performance to gain explainability, or vice versa? Here we present a comprehensive study to assess algorithm and feature influences on model performance in chemical toxicity research. We conducted over 5000 models for a Tox21 bioassay dataset of 65 assays and ~7600 compounds. Seven molecular representations as features and twelve modeling approaches varying in complexity and explainability were employed to systematically investigate the impact of various factors on model performance and explainability. We demonstrated that endpoints dictated a model’s performance, regardless of the chosen modeling approach including deep learning and chemical features. Overall, more complex models such as (LS-)SVM and Random Forest performed marginally better than simpler models such as linear regression and KNN in the presented Tox21 data analysis. Since a simpler model with acceptable performance often also is easy to interpret for the Tox21 dataset, it clearly was the preferred choice due to its better explainability. Given that each dataset had its own error structure both for dependent and independent variables, we strongly recommend that it is important to conduct a systematic study with a broad range of model complexity and feature explainability to identify model balancing its predictivity and explainability.
DOI: 10.1007/978-1-4939-6346-1_12
发表时间: 2016-01-01
期刊: HIGH-THROUGHPUT SCREENING ASSAYS IN TOXICOLOGY
影响因子: --
作者:
Huang, Ruili
通讯作者: Huang, Ruili
DOI: 10.1186/s13321-018-0258-y
发表时间: 2018-02-06
影响因子: 8.6
作者:
Moriwaki H;Tian YS;Kawashita N;Takagi T
通讯作者: Takagi T
DOI: 10.1021/acs.chemrestox.5b00481
发表时间: 2016-05-16
影响因子: 4.1
作者:
Novotarskyi S;Abdelaziz A;Sushko Y;Körner R;Vogt J;Tetko IV
通讯作者: Tetko IV
DOI: 10.3389/fenvs.2015.00080
发表时间: 2016-01-01
影响因子: 4.6
作者:
Mayr, Andreas;Klambauer, Gunter;Hochreiter, Sepp
通讯作者: Hochreiter, Sepp
DOI: 10.1038/ncomms10425
发表时间: 2016-01-26
影响因子: 16.6
作者:
Huang R;Xia M;Sakamuru S;Zhao J;Shahane SA;Attene-Ramos M;Zhao T;Austin CP;Simeonov A
通讯作者: Simeonov A