Integrative analysis and variable selection with multiple high-dimensional data sets

Integrative analysis and variable selection with multiple high-dimensional data sets
复制标题

DOI:
10.1093/biostatistics/kxr004
复制
发表时间:
2011-10-01
期刊:
影响因子:
2.1
通讯作者:
Song, Xiao
Song, Xiao
中科院分区:
数学2区
文献类型:
--
作者:
Ma, Shuangge;Huang, Jian;Song, Xiao

文献摘要

被引文献

相似文献

在高通量组学研究中,由于样本的限制,从单个数据集的分析中鉴定的标记物通常缺乏重现性。一个具有成本效益的补救办法是汇集多项可比研究的数据,并进行综合分析。由于数据的高维性和研究间的异质性,多组学数据集的整合分析具有挑战性。在这篇文章中,标记的选择,从多个异质性研究的数据进行综合分析,我们提出了一个2-范数组桥惩罚的方法。这种方法可以有效地识别多项研究中具有一致效应的标志物,并适应研究之间的异质性。我们提出了一个有效的计算算法,并建立了渐近一致性。在癌症分析研究中的模拟和应用表明,所提出的方法令人满意的性能。
In high-throughput -omics studies, markers identified from analysis of single data sets often suffer from a lack of reproducibility because of sample limitation. A cost-effective remedy is to pool data from multiple comparable studies and conduct integrative analysis. Integrative analysis of multiple -omics data sets is challenging because of the high dimensionality of data and heterogeneity among studies. In this article, for marker selection in integrative analysis of data from multiple heterogeneous studies, we propose a 2-norm group bridge penalization approach. This approach can effectively identify markers with consistent effects across multiple studies and accommodate the heterogeneity among studies. We propose an efficient computational algorithm and establish the asymptotic consistency property. Simulations and applications in cancer profiling studies show satisfactory performance of the proposed approach.