KekuleScope: prediction of cancer cell line sensitivity and compound potency using convolutional neural networks trained on compound images

KekuleScope: prediction of cancer cell line sensitivity and compound potency using convolutional neural networks trained on compound images
复制标题

DOI:
10.1186/s13321-019-0364-5
复制
发表时间:
2019-06-19
影响因子:
8.6
通讯作者:
Bender, Andreas
Bender, Andreas
中科院分区:
化学2区
文献类型:
--
作者:
Cortes-Ciriano, Isidro;Bender, Andreas

文献摘要

被引文献

相似文献

卷积神经网络(ConvNets)用于处理高含量的筛选图像或2D化合物表示在药物发现中正受到越来越多的关注。然而,现有的应用程序通常需要用于训练的大数据集,或者复杂的预训练方案。在这里,我们使用来自ChEMBL23的33个IC50数据集表明,通过扩展现有的结构(AlexNet、DenseNet-201、ResNet152和VGG-19),可以仅从它们的Kekule结构表示在连续的尺度上准确地预测化合物对癌细胞和蛋白质靶标的体外活性,这些结构是在不相关的图像数据集上预先训练的。我们证明了生成的模型的预测能力与随机森林(RF)模型和基于圆形(Morgan)指纹的全连接深度神经网络的预测能力相当,只需要标准的2D复合表示作为输入。值得注意的是,包括额外的完全连接层进一步提高了ConvNet的预测能力,最高可达10%。对RF模型和ConvNet生成的预测的分析表明,通过简单地对RF模型和ConvNet的输出进行平均,我们在多个数据集的预测中获得了比单独使用这两个模型获得的预测误差小得多的误差,这表明由ConvNet的卷积层提取的特征为Morgan指纹提供了互补的预测信号。最后,我们证明了在复合图像上训练的多任务ConvNets允许在连续的尺度上模拟COX异构体的选择性,预测的误差与数据的不确定性相当。总之,在这项工作中,我们提出了一组ConvNet体系结构,用于从它们的Kekule结构表示预测复合活动,具有最先进的性能,不需要生成复合描述符或使用复杂的图像处理技术。Https://github.com/isidroc/kekulescope.提供了复制本研究结果和所有数据集所需的代码
The application of convolutional neural networks (ConvNets) to harness high-content screening images or 2D compound representations is gaining increasing attention in drug discovery. However, existing applications often require large data sets for training, or sophisticated pretraining schemes. Here, we show using 33 IC50 data sets from ChEMBL 23 that the in vitro activity of compounds on cancer cell lines and protein targets can be accurately predicted on a continuous scale from their Kekule structure representations alone by extending existing architectures (AlexNet, DenseNet-201, ResNet152 and VGG-19), which were pretrained on unrelated image data sets. We show that the predictive power of the generated models, which just require standard 2D compound representations as input, is comparable to that of Random Forest (RF) models and fully-connected Deep Neural Networks trained on circular (Morgan) fingerprints. Notably, including additional fully-connected layers further increases the predictive power of the ConvNets by up to 10%. Analysis of the predictions generated by RF models and ConvNets shows that by simply averaging the output of the RF models and ConvNets we obtain significantly lower errors in prediction for multiple data sets, although the effect size is small, than those obtained with either model alone, indicating that the features extracted by the convolutional layers of the ConvNets provide complementary predictive signal to Morgan fingerprints. Lastly, we show that multi-task ConvNets trained on compound images permit to model COX isoform selectivity on a continuous scale with errors in prediction comparable to the uncertainty of the data. Overall, in this work we present a set of ConvNet architectures for the prediction of compound activity from their Kekule structure representations with state-of-the-art performance, that require no generation of compound descriptors or use of sophisticated image processing techniques. The code needed to reproduce the results presented in this study and all the data sets are provided at https://github.com/isidroc/kekulescope.