Towards Better BBB Passage Prediction Using an Extensive and Curated Data Set

Towards Better BBB Passage Prediction Using an Extensive and Curated Data Set
复制标题

DOI:
10.1002/minf.201400118
复制
发表时间:
2015-05-01
影响因子:
3.6
通讯作者:
Cherkasov, Artem
Cherkasov, Artem
中科院分区:
医学4区
文献类型:
--
作者:
Brito-Sanchez, Yoan;Marrero-Ponce, Yovani;Cherkasov, Artem

文献摘要

被引文献

相似文献

在本报告中,具有挑战性的任务,通过计算的方法来解决药物输送的血脑屏障(BBB)。使用分类和回归方案对一个新的广泛和精心策划的数据集(据我们所知是最大的)的对数BB进行BBB通道建模。在模型开发之前,进行了数据分析步骤,包括化学数据管理、结构、截止和聚类分析(CA)。线性判别分析(LDA)和多元线性回归(MLR)被用来拟合分类和相关函数。最好的基于LDA的模型在训练集和测试集上的总体准确率分别超过85%和83%。还开发了一个基于MLR的模型,其对实验对数BB中的方差的解释超过69%。一个简短的和一般的解释建议的模型允许的估计如何'接近'我们的计算方法是决定通过BBB的分子的因素。在最后的努力中,考虑了一些流行的和强大的机器学习方法。相对于更简单的线性技术,观察到相当或相似的性能。将大多数具有异常行为的化合物放在一个有争议的集合中,并对这些化合物进行了讨论。最后,我们的研究结果进行了比较,与以前报道的方法在文献中显示可比更好的结果。这些结果可以代表所有科学界在神经药物发现/开发项目的早期阶段可用和可重复的有用工具。
In the present report, the challenging task of drug delivery across the blood-brain barrier (BBB) is addressed via a computational approach. The BBB passage was modeled using classification and regression schemes on a novel extensive and curated data set (the largest to the best of our knowledge) in terms of log BB. Prior to the model development, steps of data analysis that comprise chemical data curation, structural, cutoff and cluster analysis (CA) were conducted. Linear Discriminant Analysis (LDA) and Multiple Linear Regression (MLR) were used to fit classification and correlation functions. The best LDA-based model showed overall accuracies over 85% and 83% for the training and test sets, respectively. Also a MLR-based model with acceptable explanation of more than 69% of the variance in the experimental log BB was developed. A brief and general interpretation of proposed models allowed the estimation on how 'near' our computational approach is to the factors that determine the passage of molecules through the BBB. In a final effort some popular and powerful Machine Learning methods were considered. Comparable or similar performance was observed respect to the simpler linear techniques. Most of the compounds with anomalous behavior were put aside into a set denoted as controversial set and discussion regarding to these compounds is provided. Finally, our results were compared with methodologies previously reported in the literature showing comparable to better results. The results could represent useful tools available and reproducible by all scientific community in the early stages of neuropharmaceutical drug discovery/development projects.