One- to four-dimensional kernels for virtual screening and the prediction of physical, chemical, and biological properties

One- to four-dimensional kernels for virtual screening and the prediction of physical, chemical, and biological properties
复制标题

DOI:
10.1021/ci600397p
复制
发表时间:
2007-05-01
影响因子:
5.6
通讯作者:
Baldi, Pierre
Baldi, Pierre
中科院分区:
化学2区
文献类型:
--
作者:
Azencott, Chloe-Agathe;Ksikes, Alexandre;Baldi, Pierre

文献摘要

被引文献

相似文献

许多化学信息学应用,包括高通量虚拟筛选,受益于能够快速预测小分子的物理,化学和生物学特性,以筛选大型储存库并识别合适的候选物。当训练集可用时,机器学习方法为这些预测提供了从头算方法的有效替代方案。在这里,我们利用丰富的分子表示,包括1D SMILES字符串,2D键图和3D坐标,以获得有效的机器学习内核来解决回归问题。我们进一步扩展了可用的光谱内核的小分子开发的分类问题,包括2.5D表面和3D内核使用Delaunay四面体化和其他技术从计算几何,3D药效团内核,3.5D或4D内核能够考虑到多个分子的配置,如构象。使用交叉验证和冗余减少方法对回归问题进行全面测试,使用几个可用的数据集来预测沸点、熔点、水溶解度、辛醇/水分配系数和生物活性,并获得最先进的结果。当有足够的训练数据可用时,2D光谱内核通常倾向于产生最好和最稳健的结果,优于现有技术。在包含数千个分子的数据集上,内核实现了水溶性预测的平方相关系数为0.91,辛醇/水分配系数预测的平方相关系数为0.94。对构象进行平均可以提高基于分子三维结构的内核的性能,特别是在具有挑战性的数据集上。水溶性(kSOL)、LogP(kLOGP)和熔点(kMELT)的核心预测因子可通过以下网址获得:http://cdb.ics.uci.edu。
Many chemoinformatics applications, including high-throughput virtual screening, benefit from being able to rapidly predict the physical, chemical, and biological properties of small molecules to screen large repositories and identify suitable candidates. When training sets are available, machine learning methods provide an effective alternative to ab initio methods for these predictions. Here, we leverage rich molecular representations including 1D SMILES strings, 2D graphs of bonds, and 3D coordinates to derive efficient machine learning kernels to address regression problems. We further expand the library of available spectral kernels for small molecules developed for classification problems to include 2.5D surface and 3D kernels using Delaunay tetrahedrization and other techniques from computational geometry, 3D pharmacophore kernels, and 3.5D or 4D kernels capable of taking into account multiple molecular configurations, such as conformers. The kernels are comprehensively tested using cross-validation and redundancy-reduction methods on regression problems using several available data sets to predict boiling points, melting points, aqueous solubility, octanol/water partition coefficients, and biological activity with state-of-the art results. When sufficient training data are available, 2D spectral kernels in general tend to yield the best and most robust results, better than state-of-the art. On data sets containing thousands of molecules, the kernels achieve a squared correlation coefficient of 0.91 for aqueous solubility prediction and 0.94 for octanol/water partition coefficient prediction. Averaging over conformations improves the performance of kernels based on the three-dimensional structure of molecules, especially on challenging data sets. Kernel predictors for aqueous solubility (kSOL), LogP (kLOGP), and melting point (kMELT) are available over the Web through: http:// cdb.ics.uci.edu.