Computational prediction of plasma protein binding of cyclic peptides from small molecule experimental data using sparse modeling techniques

Computational prediction of plasma protein binding of cyclic peptides from small molecule experimental data using sparse modeling techniques
复制标题

DOI:
10.1186/s12859-018-2529-z
复制
发表时间:
2018-12-31
期刊:
影响因子:
3
通讯作者:
Akiyama, Yutaka
Akiyama, Yutaka
中科院分区:
生物学4区
文献类型:
--
作者:
Tajimi, Takashi;Wakui, Naoki;Akiyama, Yutaka

文献摘要

被引文献

相似文献

背景基于环肽的药物发现由于其避免靶蛋白耗竭的潜力而吸引了越来越多的兴趣。在药物发现中,将药物的生物稳定性维持在适当的范围内非常重要。血浆蛋白结合(PPB)是生物稳定性最重要的指标,开发预测药物候选化合物PPB的计算方法有助于加速药物发现研究。迄今为止,已经利用机器学习对小分子药物化合物进行了 PPB 预测;然而,由于环肽的实验信息匮乏,目前还没有研究对环肽进行研究。结果首先,我们采用稀疏建模和小分子信息构建了环肽的PPB预测模型。由于环肽数据有限,应用多维非线性模型涉及过度拟合的问题。然而,稀疏建模构建的模型可以避免过度拟合,提供较高的泛化性能和可解释性。可以获得超过 1000 个小分子的 PPB 数据,我们使用它们构建了具有两种枚举方法的预测模型:枚举套索解(ELS)和前向束搜索(FBS)。在小分子化合物数据集的交叉验证上,ELS和FBS构建的预测模型的准确性等于或优于传统非线性模型(MAE=0.167-0.174)。此外,我们表明环肽的预测精度接近于小分子化合物的预测精度(MAE=0.194-0.288)。如此高的准确度无法通过直接通过套索回归(MAE=0.286-0.671)或岭回归(MAE=0.244-0.354)从环肽数据学习的简单方法获得。结论在本研究中,我们提出了一种机器学习技术,使用低维稀疏模型来计算预测环肽的PPB值。低维稀疏模型不仅表现出优异的泛化性能,而且提高了对预测模型的解释。这可以为未来的环肽药物发现研究提供共同的、值得注意的知识。
BackgroundCyclic peptide-based drug discovery is attracting increasing interest owing to its potential to avoid target protein depletion. In drug discovery, it is important to maintain the biostability of a drug within the proper range. Plasma protein binding (PPB) is the most important index of biostability, and developing a computational method to predict PPB of drug candidate compounds contributes to the acceleration of drug discovery research. PPB prediction of small molecule drug compounds using machine learning has been conducted thus far; however, no study has investigated cyclic peptides because experimental information of cyclic peptides is scarce.ResultsFirst, we adopted sparse modeling and small molecule information to construct a PPB prediction model for cyclic peptides. As cyclic peptide data are limited, applying multidimensional nonlinear models involves concerns regarding overfitting. However, models constructed by sparse modeling can avoid overfitting, offering high generalization performance and interpretability. More than 1000 PPB data of small molecules are available, and we used them to construct a prediction models with two enumeration methods: enumerating lasso solutions (ELS) and forward beam search (FBS). The accuracies of the prediction models constructed by ELS and FBS were equal to or better than those of conventional non-linear models (MAE=0.167-0.174) on cross-validation of a small molecule compound dataset. Moreover, we showed that the prediction accuracies for cyclic peptides were close to those for small molecule compounds (MAE=0.194-0.288). Such high accuracy could not be obtained by a simple method of learning from cyclic peptide data directly by lasso regression (MAE=0.286-0.671) or ridge regression (MAE=0.244-0.354).ConclusionIn this study, we proposed a machine learning techniques that uses low-dimensional sparse modeling to predict the PPB value of cyclic peptides computationally. The low-dimensional sparse model not only exhibits excellent generalization performance but also improves interpretation of the prediction model. This can provide common an noteworthy knowledge for future cyclic peptide drug discovery studies.