Averaged gene expressions for regression

Averaged gene expressions for regression
复制标题

DOI:
10.1093/biostatistics/kxl002
复制
发表时间:
2007-04-01
期刊:
影响因子:
2.1
通讯作者:
Tibshirani, Robert
Tibshirani, Robert
中科院分区:
数学2区
文献类型:
--
作者:
Park, Mee Young;Hastie, Trevor;Tibshirani, Robert

文献摘要

被引文献

相似文献

虽然平均是一种简单的技术,但它在减少方差方面发挥着重要作用。我们在DNA微阵列数据的回归中使用了平均的这一基本属性,这带来了具有比样本多得多的特征的挑战。在本文中,我们介绍了一种结合(1)层次聚类和(2)套索的两步法。通过对层次聚类得到的聚类内的基因进行平均,我们定义了超基因,并用它们来拟合回归模型,从而获得了简洁的解释和准确的结果。我们的方法得到了理论证明的支持,并在模拟和真实数据集上进行了演示。
Although averaging is a simple technique, it plays an important role in reducing variance. We use this essential property of averaging in regression of the DNA microarray data, which poses the challenge of having far more features than samples. In this paper, we introduce a two-step procedure that combines (1) hierarchical clustering and (2) Lasso. By averaging the genes within the clusters obtained from hierarchical clustering, we define supergenes and use them to fit regression models, thereby attaining concise interpretation and accuracy. Our methods are supported with theoretical justifications and demonstrated on simulated and real data sets.