Gene network-based cancer prognosis analysis with sparse boosting.

Gene network-based cancer prognosis analysis with sparse boosting.
复制标题

DOI:
10.1017/s0016672312000419
复制
发表时间:
2012-08
期刊:
影响因子:
1.5
通讯作者:
Fang, Kuangnan
Fang, Kuangnan
中科院分区:
生物学4区
文献类型:
--
作者:
Ma, Shuangge;Huang, Yuan;Huang, Jian;Fang, Kuangnan

文献摘要

被引文献

相似文献

高通量基因分析研究已经广泛进行,寻找与癌症发展和进展相关的标志物。在这项研究中,我们分析了癌症预后的研究与右删失的生存反应。利用基因表达数据,采用加权基因共表达网络分析(WGCNA)来描述基因间的相互作用。在网络分析中,节点代表基因。节点的子集称为模块,它们彼此紧密连接。相同模块内的基因往往具有共同调节的生物学功能。对于具有基因表达测量的癌症预后数据,我们的目标是识别癌症标志物,同时适当考虑网络模块结构。提出了一种两步稀疏提升方法,称为网络稀疏提升(NSBoost),用于标记选择。在第一步中,对于每个模块单独地,我们使用稀疏增强方法进行模块内标记选择并构建模块级的“超级标记”。在第二步中,我们使用超级标记来表示相同模块中所有基因的影响,并使用稀疏增强方法进行模块级选择。模拟研究表明,NSBoost可以更准确地识别癌症相关基因和模块比替代品。在乳腺癌和淋巴瘤预后研究的分析中,NSBoost确定了具有重要生物学意义的基因。它通过识别更少数量的基因/模块和/或具有更好的预测性能而优于包括增强和惩罚方法的替代方案。
High-throughput gene profiling studies have been extensively conducted, searching for markers associated with cancer development and progression. In this study, we analyse cancer prognosis studies with right censored survival responses. With gene expression data, we adopt the weighted gene co-expression network analysis (WGCNA) to describe the interplay among genes. In network analysis, nodes represent genes. There are subsets of nodes, called modules, which are tightly connected to each other. Genes within the same modules tend to have co-regulated biological functions. For cancer prognosis data with gene expression measurements, our goal is to identify cancer markers, while properly accounting for the network module structure. A two-step sparse boosting approach, called Network Sparse Boosting (NSBoost), is proposed for marker selection. In the first step, for each module separately, we use a sparse boosting approach for within-module marker selection and construct module-level ‘super markers ’. In the second step, we use the super markers to represent the effects of all genes within the same modules and conduct module-level selection using a sparse boosting approach. Simulation study shows that NSBoost can more accurately identify cancer-associated genes and modules than alternatives. In the analysis of breast cancer and lymphoma prognosis studies, NSBoost identifies genes with important biological implications. It outperforms alternatives including the boosting and penalization approaches by identifying a smaller number of genes/modules and/or having better prediction performance.