Benchmarking Small-Dataset Structure-Activity-Relationship Models for Prediction of Wnt Signaling Inhibition

Benchmarking Small-Dataset Structure-Activity-Relationship Models for Prediction of Wnt Signaling Inhibition
复制标题

DOI:
10.1109/access.2020.3046190
复制
发表时间:
2020-01-01
期刊:
影响因子:
3.9
通讯作者:
Xu, Guangyu
Xu, Guangyu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kokabi, Mahtab;Donnelly, Matthew;Xu, Guangyu

文献摘要

被引文献

相似文献

基于机器学习算法的定量构效关系(QSAR)模型是加速药物发现过程和治疗发展的有力工具。考虑到获取大型训练数据集的成本,检查QSAR分析是否可以合理地预测药物活性是有用的,只使用一个小数据集(大小为100),并在特定应用的研究中对这些小数据集QSAR模型进行基准测试。为此,我们在这里提出了一项针对小数据集QSAR模型的系统基准研究,该模型是为预测有效的Wnt信号抑制物而建立的,这些抑制物对人类流行疾病(如癌症)的治疗开发至关重要。具体地说,我们基于4个性能最好的算法、6个常用的分子指纹和3个典型的指纹长度,检查了总共72个二维(2D)QSAR模型。我们使用训练数据集(56个化合物)训练这些模型,在4个品质因数(FOM)上对它们的性能进行基准测试,并使用外部验证数据集(14个化合物)检查它们的预测准确性。我们的数据表明,当:1)选择分子指纹来提供充分、唯一且不太详细的药物化合物化学结构的表示时,模型的性能是最大的;2)选择算法来减少由于数据集中类别失衡而导致的错误预测的数量;以及3)选择模型来在所有4个FOM上达到平衡的性能。这些结果可能为开发用于药物活性预测的高性能小数据集QSAR模型提供一般指导。
Quantitative structure-activity relationship (QSAR) models based on machine learning algorithms are powerful tools to expedite drug discovery processes and therapeutics development. Given the cost in acquiring large-sized training datasets, it is useful to examine if QSAR analysis can reasonably predict drug activity with only a small-sized dataset (size < 100) and benchmark these small-dataset QSAR models in application-specific studies. To this end, here we present a systematic benchmarking study on small-dataset QSAR models built for prediction of effective Wnt signaling inhibitors, which are essential to therapeutics development in prevalent human diseases (e.g., cancer). Specifically, we examined a total of 72 two-dimensional (2D) QSAR models based on 4 best-performing algorithms, 6 commonly used molecular fingerprints, and 3 typical fingerprint lengths. We trained these models using a training dataset (56 compounds), benchmarked their performance on 4 figures-of-merit (FOMs), and examined their prediction accuracy using an external validation dataset (14 compounds). Our data show that the model performance is maximized when: 1) molecular fingerprints are selected to provide sufficient, unique, and not overly detailed representations of the chemical structures of drug compounds; 2) algorithms are selected to reduce the number of false predictions due to class imbalance in the dataset; and 3) models are selected to reach balanced performance on all 4 FOMs. These results may provide general guidelines in developing high-performance small-dataset QSAR models for drug activity prediction.