Dual-dropout graph convolutional network for predicting synthetic lethality in human cancers

Dual-dropout graph convolutional network for predicting synthetic lethality in human cancers
复制标题

用于预测人类癌症综合致死率的双丢失图卷积网络

DOI:
10.1093/bioinformatics/btaa211
复制
发表时间:
2020-08-15
期刊:
影响因子:
5.8
通讯作者:
Hao, Yuexing
Hao, Yuexing
中科院分区:
生物学3区
文献类型:
--
作者:
Cai, Ruichu;Chen, Xuexin;Hao, Yuexing

文献摘要

被引文献

相似文献

动机:合成致死(SL)是一种有前途的癌症治疗基因相互作用形式,因为它能够识别靶向癌细胞的特定基因,而不破坏正常细胞。由于高通量湿实验室设置通常成本高昂并面临各种挑战,计算方法已成为一种实用的补充。特别地,预测SL可以被公式化为相互作用基因的图上的链接预测任务。虽然矩阵分解技术已被广泛采用的链接预测,他们专注于映射基因的潜在表示在隔离,没有聚集来自相邻基因的信息。图卷积网络(GCN)可以捕获图中的这种邻域依赖性。然而,它仍然是具有挑战性的应用GCN SL预测SL的相互作用是非常稀疏的,这是更有可能导致overfitting.Results:在这篇文章中,我们提出了一种新的双辍学GCN(DDGCN)学习更强大的基因表示SL预测。我们采用粗粒度的节点丢弃和细粒度的边缘丢弃来解决这个问题,标准丢弃在香草GCN通常是不足以减少稀疏图上的过拟合。特别是,粗粒度节点丢弃可以有效地和系统地在节点(基因)级别强制丢弃,而细粒度边缘丢弃可以进一步微调交互(边缘)级别的丢弃。我们进一步提出了一个理论框架来证明我们的模型架构。最后,我们在人类SL数据集上进行了广泛的实验,结果表明,与最先进的方法相比,我们的模型具有上级性能。
Motivation: Synthetic lethality (SL) is a promising form of gene interaction for cancer therapy, as it is able to identify specific genes to target at cancer cells without disrupting normal cells. As high-throughput wet-lab settings are often costly and face various challenges, computational approaches have become a practical complement. In particular, predicting SLs can be formulated as a link prediction task on a graph of interacting genes. Although matrix factorization techniques have been widely adopted in link prediction, they focus on mapping genes to latent representations in isolation, without aggregating information from neighboring genes. Graph convolutional networks (GCN) can capture such neighborhood dependency in a graph. However, it is still challenging to apply GCN for SL prediction as SL interactions are extremely sparse, which is more likely to cause overfitting.Results: In this article, we propose a novel dual-dropout GCN (DDGCN) for learning more robust gene representations for SL prediction. We employ both coarse-grained node dropout and fine-grained edge dropout to address the issue that standard dropout in vanilla GCN is often inadequate in reducing overfitting on sparse graphs. In particular, coarse-grained node dropout can efficiently and systematically enforce dropout at the node (gene) level, while fine-grained edge dropout can further fine-tune the dropout at the interaction (edge) level. We further present a theoretical framework to justify our model architecture. Finally, we conduct extensive experiments on human SL datasets and the results demonstrate the superior performance of our model in comparison with state-of-the-art methods.