Multiple hypothesis testing by clustering treatment effects

Multiple hypothesis testing by clustering treatment effects
复制标题

DOI:
10.1198/016214507000000211
复制
发表时间:
2007-06-01
影响因子:
3.7
通讯作者:
Newton, Michael A.
Newton, Michael A.
中科院分区:
数学1区
文献类型:
--
作者:
Dahl, David B.;Newton, Michael A.

文献摘要

被引文献

相似文献

多重假设检验和聚类一直是高维推理中广泛研究的主题,但这些问题通常被单独处理。通过根据共享参数值定义真正的集群,我们可以提高单个测试的灵敏度,因为可以获得更多与相同参数值相关的数据。我们开发并评估了一种混合方法,该方法使用聚类信息来提高测试灵敏度并适应真实聚类中的不确定性。为了研究混合方法的潜在功效,我们首先研究一个程式化的示例,其中每个对象都使用标准 z 分数进行评估,但不同的对象通过共享参数值连接。我们表明,当聚类估计得足够好时,测试能力就会增加。接下来,我们使用共轭狄利克雷过程混合模型开发基于模型的分析。该方法是通用的,但为了特异性,我们将注意力集中在微阵列基因表达数据上,并积极应用聚类和多种测试方法。聚类提供了在基因之间共享信息的方法,混合方法通过马尔可夫链抽样对这些聚类中的不确定性进行平均。仿真表明,当聚类程度较高或中等时,混合方法的性能明显优于其他方法,即使在弱聚类情况下也能表现良好。所提出的方法通过来自衰老对心脏组织基因表达影响的研究的微阵列数据进行了说明。
Multiple hypothesis testing and clustering have been the subject of extensive research in high-dimensional inference, yet these problems usually have been treated separately. By defining true clusters in terms of shared parameter values, we could improve the sensitivity of individual tests, because more data bearing on the same parameter values are available. We develop and evaluate a hybrid methodology that uses clustering information to increase testing sensitivity and accommodates uncertainty in the true clustering. To investigate the potential efficacy of the hybrid approach, we first study a stylized example in which each object is evaluated with a standard z score but different objects are connected by shared parameter values. We show that there is increased testing power when the clustering is estimated sufficiently well. We next develop a model-based analysis using a conjugate Dirichlet process mixture model. The method is' general, but for specificity we focus attention on microarray gene expression data, to which both clustering and multiple testing methods are actively applied. Clusters provide the means for sharing information among genes, and the hybrid methodology averages over uncertainty in these clusters through Markov chain sampling. Simulations show that the hybrid method performs substantially better than other methods when clustering is heavy or moderate and performs well even under weak clustering. The proposed method is illustrated on microarray data from a study of the effects of aging on gene expression in heart tissue.