Group additive regression models for genomic data analysis

Group additive regression models for genomic data analysis
复制标题

用于基因组数据分析的组加性回归模型

DOI:
10.1093/biostatistics/kxm015
复制
发表时间:
2008-01-01
期刊:
影响因子:
2.1
通讯作者:
Li, Hongzhe
Li, Hongzhe
中科院分区:
数学2区
文献类型:
--
作者:
Luan, Yihui;Li, Hongzhe

文献摘要

被引文献

相似文献

基因组研究中的一个重要问题是识别与临床表型相关的基因组特征,如基因表达数据或DNA单核苷酸多态(SNPs)。通常,这些基因组数据可以自然地分成具有生物学意义的组,如属于相同途径的基因或基因内的SNP。在这篇文章中,我们提出了群体加性回归模型和群体梯度下降提升程序来识别与临床表型相关的基因组特征组。我们的模拟结果表明,通过将变量划分到适当的组中,我们可以更好地识别与表型相关的组特征。此外,预测均方误差也比分量升压法小。我们展示了这些方法在乳腺癌微阵列基因表达数据的基于路径的分析中的应用。对乳腺癌微阵列基因表达数据集的分析结果表明,金属内肽酶(MMPs)和基质金属蛋白酶抑制物(MMPs)的途径,以及细胞的增殖、生长和维持对乳腺癌特异性生存至关重要。
One important problem in genomic research is to identify genomic features such as gene expression data or DNA single nucleotide polymorphisms (SNPs) that are related to clinical phenotypes. Often these genomic data can be naturally divided into biologically meaningful groups such as genes belonging to the same pathways or SNPs within genes. In this paper, we propose group additive regression models and a group gradient descent boosting procedure for identifying groups of genomic features that are related to clinical phenotypes. Our simulation results show that by dividing the variables into appropriate groups, we can obtain better identification of the group features that are related to the phenotypes. In addition, the prediction mean square errors are also smaller than the component-wise boosting procedure. We demonstrate the application of the methods to pathway-based analysis of microarray gene expression data of breast cancer. Results from analysis of a breast cancer microarray gene expression data set indicate that the pathways of metalloendopeptidases (MMPs) and MMP inhibitors, as well as cell proliferation, cell growth, and maintenance are important to breast cancer-specific survival.