Integrate multi-omics data with biological interaction networks using Multi-view Factorization AutoEncoder (MAE)

Integrate multi-omics data with biological interaction networks using Multi-view Factorization AutoEncoder (MAE)
复制标题

利用多视图分解自动编码器(MAE)集成多组学数据与生物相互作用网络

DOI:
10.1186/s12864-019-6285-x
复制
发表时间:
2019-12-20
期刊:
影响因子:
4.4
通讯作者:
Zhang, Aidong
Zhang, Aidong
中科院分区:
生物学2区
文献类型:
--
作者:
Ma, Tianle;Zhang, Aidong

文献摘要

被引文献

相似文献

背景:各种癌症和其他疾病的全面分子谱分析产生了大量的多组学数据。每种类型的组学数据对应一个特征空间,如基因表达,miRNA表达,DNA甲基化等。整合多组学数据可以连接不同层的分子特征空间,对于阐明各种疾病的分子途径至关重要。挖掘多组学数据的机器学习方法在揭示分子特征之间的复杂关系方面具有很大的潜力。然而,由于“大p,小n”问题(即,结果:我们开发了一种具有网络约束的多视图分解自动编码器(Multi-view Factorization AutoEncoder,MAE)方法,可以无缝集成多组学数据和领域知识,如分子相互作用网络。我们的方法通过深度表征学习同时学习特征和患者嵌入。特征表示和患者表示都受到在训练目标中指定为正则化项的某些约束。通过将领域知识纳入训练目标,我们隐式地将良好的归纳偏差引入机器学习模型,这有助于提高模型的泛化能力。我们在TCGA数据集上进行了广泛的实验,并证明了使用我们提出的方法来预测目标临床变量的多组学数据和生物相互作用网络的集成能力。结论:为了缓解多组学数据深度学习中的过拟合问题和“大p,小n”问题,将生物领域知识作为归纳偏差纳入模型是有帮助的。这是非常有前途的设计机器学习模型,促进大规模的多组学数据和生物医学领域的知识的无缝集成,以揭示分子特征和临床特征之间的复杂关系。
Background: Comprehensive molecular profiling of various cancers and other diseases has generated vast amounts of multi-omics data. Each type of -omics data corresponds to one feature space, such as gene expression, miRNA expression, DNA methylation, etc. Integrating multi-omics data can link different layers of molecular feature spaces and is crucial to elucidate molecular pathways underlying various diseases. Machine learning approaches to mining multi-omics data hold great promises in uncovering intricate relationships among molecular features. However, due to the "big p, small n" problem (i.e., small sample sizes with high-dimensional features), training a large-scale generalizable deep learning model with multi-omics data alone is very challenging.Results: We developed a method called Multi-view Factorization AutoEncoder (MAE) with network constraints that can seamlessly integrate multi-omics data and domain knowledge such as molecular interaction networks. Our method learns feature and patient embeddings simultaneously with deep representation learning. Both feature representations and patient representations are subject to certain constraints specified as regularization terms in the training objective. By incorporating domain knowledge into the training objective, we implicitly introduced a good inductive bias into the machine learning model, which helps improve model generalizability. We performed extensive experiments on the TCGA datasets and demonstrated the power of integrating multi-omics data and biological interaction networks using our proposed method for predicting target clinical variables.Conclusions: To alleviate the overfitting problem in deep learning on multi-omics data with the "big p, small n" problem, it is helpful to incorporate biological domain knowledge into the model as inductive biases. It is very promising to design machine learning models that facilitate the seamless integration of large-scale multi-omics data and biomedical domain knowledge for uncovering intricate relationships among molecular features and clinical features.