Network-constrained regularization and variable selection for analysis of genomic data

Network-constrained regularization and variable selection for analysis of genomic data
复制标题

DOI:
10.1093/bioinformatics/btn081
复制
发表时间:
2008-05-01
期刊:
影响因子:
5.8
通讯作者:
Li, Hongzhe
Li, Hongzhe
中科院分区:
生物学3区
文献类型:
--
作者:
Li, Caiyan;Li, Hongzhe

文献摘要

被引文献

相似文献

动机:图表或网络是描述信息的常见方式。特别是在生物学中,许多不同的生物过程都用图表来表示,例如调控网络或代谢途径。这种在多年生物医学研究中收集的先验信息是对标准数值基因组数据(如微阵列基因表达数据)的有益补充。如何将已知的生物网络或图形编码的信息纳入数值数据的分析,提出了有趣的统计挑战。在这篇文章中,我们介绍了一个网络约束正则化过程的线性回归分析,以便将这些图的信息纳入到数值数据的分析中,其中网络表示为一个图及其相应的拉普拉斯矩阵。我们定义了一个网络约束的惩罚函数,惩罚的L-1-范数的系数,但鼓励网络上的系数的光滑性。结果:模拟研究表明,该方法是相当有效的识别基因和子网络,与疾病相关,并具有更高的灵敏度比常用的程序,不使用的通路结构信息。应用于一个胶质母细胞瘤微阵列基因表达数据集确定了与胶质母细胞瘤生存相关的几个京都基因和基因组百科全书(KEGG)转录途径上的几个子网络,其中许多得到了已发表文献的支持。拟议的网络-约束正则化过程有效地利用了已知的通路结构来识别相关的基因和可能是在一般回归框架中与表型相关。随着越来越多的生物网络被识别并记录在数据库中,所提出的方法应该在识别与疾病和其他生物过程相关的子网络方面找到更多的应用。
Motivation: Graphs or networks are common ways of depicting information. In biology in particular, many different biological processes are represented by graphs, such as regulatory networks or metabolic pathways. This kind of a priori information gathered over many years of biomedical research is a useful supplement to the standard numerical genomic data such as microarray gene-expression data. How to incorporate information encoded by the known biological networks or graphs into analysis of numerical data raises interesting statistical challenges. In this article, we introduce a network-constrained regularization procedure for linear regression analysis in order to incorporate the information from these graphs into an analysis of the numerical data, where the network is represented as a graph and its corresponding Laplacian matrix. We define a network-constrained penalty function that penalizes the L-1-norm of the coefficients but encourages smoothness of the coefficients on the network.Results: Simulation studies indicated that the method is quite effective in identifying genes and subnetworks that are related to disease and has higher sensitivity than the commonly used procedures that do not use the pathway structure information. Application to one glioblastoma microarray gene-expression dataset identified several subnetworks on several of the Kyoto Encyclopedia of Genes and Genomes (KEGG) transcriptional pathways that are related to survival from glioblastoma, many of which were supported by published literatures.Conclusions: The proposed network-constrained regularization procedure efficiently utilizes the known pathway structures in identifying the relevant genes and the subnetworks that might be related to phenotype in a general regression framework. As more biological networks are identified and documented in databases, the proposed method should find more applications in identifying the subnetworks that are related to diseases and other biological processes.Contact: hongzhe@mail.med.upenn.edu.