Studying and mitigating the effects of data drifts on ML model performance at the example of chemical toxicity data.

Studying and mitigating the effects of data drifts on ML model performance at the example of chemical toxicity data.
复制标题

DOI:
10.1038/s41598-022-09309-3
复制
发表时间:
2022-05-04
期刊:
影响因子:
4.6
通讯作者:
Volkamer, Andrea
Volkamer, Andrea
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Morger, Andrea;de Lomana, Marina Garcia;Norinder, Ulf;Svensson, Fredrik;Kirchmair, Johannes;Mathea, Miriam;Volkamer, Andrea

文献摘要

参考文献

被引文献

相似文献

机器学习模型被广泛应用于预测特定蛋白质上小分子的分子特性或生物活性。模型可以集成在共形预测(CP)框架中,该框架增加了校准步骤以估计预测的置信度。CP模型的优点是在假设测试集和校准集是可交换的情况下确保预定义的误差率。如果测试数据偏离训练数据的描述符空间,或者检测设置发生变化,则可能无法满足此假设,并且无法保证模型有效。在这项研究中,内部有效的CP模型的性能时,适用于较新的时间分割数据或外部数据进行了评估。详细地,时间数据漂移进行了分析的基础上,从ChEMBL数据库的12个数据集。此外,研究了在公开数据上训练的模型与应用于肝毒性和MNT体内终点的专有数据之间的差异。在大多数情况下,在模型的有效性急剧下降时,观察到的时间分割或外部(保持)测试集。为了克服模型有效性的降低,研究了用与保持集更相似的数据更新校准集的策略。更新校准集通常会提高有效性,在许多情况下将其完全恢复到其预期值。恢复的有效性是有信心地应用CP模型的首要条件。然而,有效性的提高是以降低模型效率为代价的,因为更多的预测被确定为不确定的。本研究提出了一种策略,重新校准CP模型,以减轻数据漂移的影响。更新校准集而无需重新训练模型已被证明是恢复大多数模型有效性的有用方法。
Machine learning models are widely applied to predict molecular properties or the biological activity of small molecules on a specific protein. Models can be integrated in a conformal prediction (CP) framework which adds a calibration step to estimate the confidence of the predictions. CP models present the advantage of ensuring a predefined error rate under the assumption that test and calibration set are exchangeable. In cases where the test data have drifted away from the descriptor space of the training data, or where assay setups have changed, this assumption might not be fulfilled and the models are not guaranteed to be valid. In this study, the performance of internally valid CP models when applied to either newer time-split data or to external data was evaluated. In detail, temporal data drifts were analysed based on twelve datasets from the ChEMBL database. In addition, discrepancies between models trained on publicly-available data and applied to proprietary data for the liver toxicity and MNT in vivo endpoints were investigated. In most cases, a drastic decrease in the validity of the models was observed when applied to the time-split or external (holdout) test sets. To overcome the decrease in model validity, a strategy for updating the calibration set with data more similar to the holdout set was investigated. Updating the calibration set generally improved the validity, restoring it completely to its expected value in many cases. The restored validity is the first requisite for applying the CP models with confidence. However, the increased validity comes at the cost of a decrease in model efficiency, as more predictions are identified as inconclusive. This study presents a strategy to recalibrate CP models to mitigate the effects of data drifts. Updating the calibration sets without having to retrain the model has proven to be a useful approach to restore the validity of most models.
DOI: 10.1093/nar/gkv352
发表时间: 2015-07-01
影响因子: 14.9
作者:
Davies M;Nowotka M;Papadatos G;Dedman N;Gaulton A;Atkinson F;Bellis L;Overington JP
通讯作者: Overington JP
DOI: 10.1289/ehp5580
发表时间: 2020-02-01
影响因子: 10.4
作者:
Mansouri, Kamel;Kleinstreuer, Nicole;Judson, Richard S.
通讯作者: Judson, Richard S.
DOI: 10.1186/s13321-020-00444-5
发表时间: 2020-06-05
影响因子: 8.6
作者:
Cortes-Ciriano, Isidro;Skuta, Ctibor;Svozil, Daniel
通讯作者: Svozil, Daniel
DOI: 10.3390/ijms21103585
发表时间: 2020-05-01
影响因子: 5.6
作者:
Mathai, Neann;Kirchmair, Johannes
通讯作者: Kirchmair, Johannes
DOI: 10.3389/fenvs.2015.00085
发表时间: 2016-01-01
影响因子: 4.6
作者:
Huang, Ruili;Xia, Menghang;Simeonov, Anton
通讯作者: Simeonov, Anton