Applying Machine Learning to Ultrafast Shape Recognition in Ligand-Based Virtual Screening

Applying Machine Learning to Ultrafast Shape Recognition in Ligand-Based Virtual Screening
复制标题

DOI:
10.3389/fphar.2019.01675
复制
发表时间:
2020-02-19
影响因子:
5.6
通讯作者:
Ebejer, Jean-Paul
Ebejer, Jean-Paul
中科院分区:
医学2区
文献类型:
--
作者:
Bonanno, Etienne;Ebejer, Jean-Paul

文献摘要

被引文献

相似文献

超快形状识别(USR)及其衍生物沿着是基于配体的虚拟筛选(LBVS)方法,其将关于分子形状以及其他性质的三维信息浓缩成一小组数字描述符。这些可以用于使用简单的逆曼哈顿距离度量来有效地计算分子对之间的相似性度量。在这项研究中,我们探索使用合适的机器学习技术,可以使用USR描述符进行训练,以提高潜在新线索的相似性检测。我们使用来自有用的Decoys-Enhanced目录的分子来构建基于三种不同算法的机器学习模型:高斯混合模型(GARCH),隔离森林和人工神经网络(ANN)。我们基于全分子构象模型以及仅最低能量构象(LEC)训练模型。我们还研究了我们的模型在较小的数据集上训练时的性能,以便在只有少量活性物先验已知时对虚拟筛选场景进行建模。我们的研究结果表明,与最先进的USR衍生方法ElectroShape 5D相比,Gestival的平均性能提高了430%,在富集因子方面比ElectroShape 5D的平均性能提高了940%。此外,我们证明了我们的模型能够保持其性能,在富集因子方面,随着训练数据集的大小逐渐减小,平均值在10%以内。此外,我们还证明了使用我们选择的机器学习模型进行回顾性筛选的运行时间比标准USR快,平均快10倍,包括训练所需的时间。我们的研究结果表明,机器学习技术可以显着提高虚拟筛选的性能和效率的USR家庭的方法。
Ultrafast Shape Recognition (USR), along with its derivatives, are Ligand-Based Virtual Screening (LBVS) methods that condense 3-dimensional information about molecular shape, as well as other properties, into a small set of numeric descriptors. These can be used to efficiently compute a measure of similarity between pairs of molecules using a simple inverse Manhattan Distance metric. In this study we explore the use of suitable Machine Learning techniques that can be trained using USR descriptors, so as to improve the similarity detection of potential new leads. We use molecules from the Directory for Useful Decoys-Enhanced to construct machine learning models based on three different algorithms: Gaussian Mixture Models (GMMs), Isolation Forests and Artificial Neural Networks (ANNs). We train models based on full molecule conformer models, as well as the Lowest Energy Conformations (LECs) only. We also investigate the performance of our models when trained on smaller datasets so as to model virtual screening scenarios when only a small number of actives are known a priori. Our results indicate significant performance gains over a state of the art USR-derived method, ElectroShape 5D, with GMMs obtaining a mean performance up to 430% better than that of ElectroShape 5D in terms of Enrichment Factor with a maximum improvement of up to 940%. Additionally, we demonstrate that our models are capable of maintaining their performance, in terms of enrichment factor, within 10% of the mean as the size of the training dataset is successively reduced. Furthermore, we also demonstrate that running times for retrospective screening using the machine learning models we selected are faster than standard USR, on average by a factor of 10, including the time required for training. Our results show that machine learning techniques can significantly improve the virtual screening performance and efficiency of the USR family of methods.