Molecular pathway identification using biological network-regularized logistic models.

Molecular pathway identification using biological network-regularized logistic models.
复制标题

DOI:
10.1186/1471-2164-14-s8-s7
复制
发表时间:
2013
期刊:
影响因子:
4.4
通讯作者:
Liu Z
Liu Z
中科院分区:
生物学2区
文献类型:
--
作者:
Zhang W;Wan YW;Allen GI;Pang K;Anderson ML;Liu Z

文献摘要

被引文献

相似文献

选择指示疾病的基因和途径是计算生物学的中心问题。在分析多维基因组数据时,这个问题尤其具有挑战性。许多工具,如基于L1范数的正则化及其扩展弹性网和熔融套索,已经被引入来应对这一挑战。然而,这些方法往往忽略了文献中策划的大量先验生物网络信息。我们提出使用图拉普拉斯正则化Logistic回归将生物网络集成到疾病分类和路径关联问题中。仿真研究表明,该算法的性能优于弹性网络和套索分析。该算法的实用性还被其使用由癌症基因组图谱(TCGA)联盟最近生成的大型乳腺癌数据集可靠地区分乳腺癌亚型的能力所验证。我们的方法确定的许多蛋白质-蛋白质相互作用模块进一步得到了文献中发表的证据的支持。该算法的源代码可在http://www.github.com/zhandong/Logit-Lapnet.上免费获得图拉普拉斯正则化Logistic回归是一种识别与疾病亚型相关的关键通路和模块的有效算法。随着我们对生物调控网络知识的快速扩展,这种方法将变得更加准确,并越来越适用于挖掘转录、表观基因组和其他类型的全基因组关联研究。
Selecting genes and pathways indicative of disease is a central problem in computational biology. This problem is especially challenging when parsing multi-dimensional genomic data. A number of tools, such as L1-norm based regularization and its extensions elastic net and fused lasso, have been introduced to deal with this challenge. However, these approaches tend to ignore the vast amount of a priori biological network information curated in the literature. We propose the use of graph Laplacian regularized logistic regression to integrate biological networks into disease classification and pathway association problems. Simulation studies demonstrate that the performance of the proposed algorithm is superior to elastic net and lasso analyses. Utility of this algorithm is also validated by its ability to reliably differentiate breast cancer subtypes using a large breast cancer dataset recently generated by the Cancer Genome Atlas (TCGA) consortium. Many of the protein-protein interaction modules identified by our approach are further supported by evidence published in the literature. Source code of the proposed algorithm is freely available at http://www.github.com/zhandong/Logit-Lapnet. Logistic regression with graph Laplacian regularization is an effective algorithm for identifying key pathways and modules associated with disease subtypes. With the rapid expansion of our knowledge of biological regulatory networks, this approach will become more accurate and increasingly useful for mining transcriptomic, epi-genomic, and other types of genome wide association studies.