Reliable Model Selection without Reference Values by Utilizing Model Diversity with Prediction Similarity

Reliable Model Selection without Reference Values by Utilizing Model Diversity with Prediction Similarity
复制标题

DOI:
10.1021/acs.jcim.0c01493
复制
发表时间:
2021-04
影响因子:
5.6
通讯作者:
Robert C. Spiers;J. Kalivas
Robert C. Spiers;J. Kalivas
中科院分区:
化学2区
文献类型:
--
作者:
Robert C. Spiers;J. Kalivas

文献摘要

被引文献

相似文献

如果选择了适当的模型,使用各种数据格式(例如近红外(NIR)光谱和定量结构活性关系(QSAR)数据)进行预测建模(校准或训练)可以提供重要信息。类似地,利用一般的模型选择方法,可以针对动态建模执行从原始建模条件到新条件的谱模型维护(更新)。基础建模(偏最小二乘(PLS)等)和维护过程(域自适应或转移学习等)需要选择调谐参数值以隔离可以准确预测新样品或分子的模型,例如,预测分析物浓度的PLS潜在变量的数量。无论建模任务如何,模型选择都是复杂的,并且没有可靠的协议。调整参数的选择通常只依赖于一个模型质量测量,使用预测精度评估模型偏差。在本文中开发的是一个通用的模型选择过程中使用的概念,从共识建模和QSAR活性景观。它是一种共识过滤方法,优先考虑模型多样性(MD),同时保留预测相似性(PS),并融合了一个共同的偏差-方差权衡措施。MDPS的一个显著特征是不需要交叉验证方案,因为模型是相对于预测新样品或分子来选择的,即,模型选择使用未标记的样本(没有参考值)进行主动预测。使用四个近红外数据集和QSAR数据集的MDPS模型选择的多功能性和可靠性。该研究还证实了罗生门效应,即没有一个最佳模型调整参数值可以提供准确的预测。
Predictive modeling (calibration or training) with various data formats, such as near-infrared (NIR) spectra and quantitative structure-activity relationship (QSAR) data, provides essential information if a proper model is selected. Similarly, with a general model selection approach, spectral model maintenance (updating) from original modeling conditions to new conditions can be performed for dynamic modeling. Fundamental modeling (partial least-squares (PLS) and others) and maintenance processes (domain adaptation or transfer learning and others) require selection of tuning parameter(s) values to isolate models that can accurately predict new samples or molecules, e.g., number of PLS latent variables to predict analyte concentration. Regardless of the modeling task, model selection is complex and without a reliable protocol. Tuning parameter selection typically depends on only one model quality measure assessing model bias using prediction accuracy. Developed in this paper is a generic model selection process using concepts from consensus modeling and QSAR activity landscapes. It is a consensus filtering approach that prioritizes model diversity (MD) while conserving prediction similarity (PS) fused with a common bias-variance trade-off measure. A significant feature of MDPS is that a cross-validation scheme is not needed because models are selected relative to predicting new samples or molecules, i.e., model selection uses unlabeled samples (without reference values) for active predictions. The versatility and reliability of MDPS model selection is shown using four NIR data sets and a QSAR data set. The study also substantiates the Rashomon effect where there is not one best model tuning parameter value that provides accurate predictions.