SAMPL6 challenge results from [Formula: see text] predictions based on a general Gaussian process model.

SAMPL6 challenge results from [Formula: see text] predictions based on a general Gaussian process model.
复制标题

DOI:
10.1007/s10822-018-0169-z
复制
发表时间:
2018-10
影响因子:
3.5
通讯作者:
Skillman AG
Skillman AG
中科院分区:
生物学3区
文献类型:
--
作者:
Bannan CC;Mobley DL;Skillman AG

文献摘要

参考文献

被引文献

相似文献

各种领域将受益于准确的pKa预测,特别是药物设计,因为电离状态的变化会对分子的理化性质产生影响。最近SAMPL 6盲法挑战的参与者被要求提交24种药物样小分子的微观和宏观pKas预测。我们最近建立了一个通用模型,用于预测PKAS使用高斯过程回归训练使用每个电离基团的物理和化学特征。我们的管道采用分子图,并使用OpenEye Toolkits来计算描述质子去除的特征。这些特征被输入Scikit-learn高斯过程以预测微观pKas,然后用于分析确定宏观pKas。我们的高斯过程是在一组2,700宏观pKa上训练的,这些pKa来自单质子和选择的双质子分子。在这里,我们分享了我们在SAMPL 6挑战中的微观和宏观预测结果。总的来说,与其他参与者相比,我们排在中间,但考虑到挑战分子的化学多样性并且通常是多质子的,而我们的训练集主要是单质子的,我们与实验的相当好的一致性仍然是有希望的。在建立这个模型时,对我们来说特别重要的是包括基于分子化学的不确定性估计,这将反映我们预测的可能准确性。我们的模型报告了似乎具有我们的适用范围之外的化学性质的分子的大的不确定性,沿着在分位数-分位数图中具有良好的一致性,表明它可以预测自己的准确性。这个挑战强调了改进我们模型的各种方法,包括在我们的训练集中添加更多的多质子分子,以及更仔细地考虑我们确定或不确定哪些官能团是可电离的。
A variety of fields would benefit from accurate pKa predictions, especially drug design due to the effect a change in ionization state can have on a molecule’s physiochemical properties. Participants in the recent SAMPL6 blind challenge were asked to submit predictions for microscopic and macroscopic pKas of 24 drug like small molecules. We recently built a general model for predicting pKas using a Gaussian process regression trained using physical and chemical features of each ionizable group. Our pipeline takes a molecular graph and uses the OpenEye Toolkits to calculate features describing the removal of a proton. These features are fed into a Scikit-learn Gaussian process to predict microscopic pKas which are then used to analytically determine macroscopic pKas. Our Gaussian process is trained on a set of 2,700 macroscopic pKas from monoprotic and select diprotic molecules. Here, we share our results for microscopic and macroscopic predictions in the SAMPL6 challenge. Overall, we ranked in the middle of the pack compared to other participants, but our fairly good agreement with experiment is still promising considering the challenge molecules are chemically diverse and often polyprotic while our training set is predominately monoprotic. Of particular importance to us when building this model was to include an uncertainty estimate based on the chemistry of the molecule that would reflect the likely accuracy of our prediction. Our model reports large uncertainties for the molecules that appear to have chemistry outside our domain of applicability, along with good agreement in quantile-quantile plots, indicating it can predict its own accuracy. The challenge highlighted a variety of means to improve our model, including adding more polyprotic molecules to our training set and more carefully considering what functional groups we do or do not identify as ionizable.
DOI: 10.1107/s0021889883010985
发表时间: 1983-01-01
影响因子: 6.1
作者:
CONNOLLY, ML
通讯作者: CONNOLLY, ML
DOI: 10.1016/s0045-6535(98)00172-6
发表时间: 1999-01-01
期刊: CHEMOSPHERE
影响因子: 8.8
作者:
Citra, MJ
通讯作者: Citra, MJ
DOI: 10.1021/ed063p246
发表时间: 1986-03-01
影响因子: 3
作者:
BODNER, GM
通讯作者: BODNER, GM
DOI: 10.1002/jcc.1032
发表时间: 2001-04-30
影响因子: 3
作者:
Grant, JA;Pickup, BT;Nicholls, A
通讯作者: Nicholls, A
DOI: 10.1002/qua.24481
发表时间: 2013-09-15
影响因子: 2.2
作者:
Bochevarov, Art D.;Harder, Edward;Friesner, Richard A.
通讯作者: Friesner, Richard A.