Multiple Testing for Pattern Identification, With Applications to Microarray Time-Course Experiments

Multiple Testing for Pattern Identification, With Applications to Microarray Time-Course Experiments
复制标题

DOI:
10.1198/jasa.2011.ap09587
复制
发表时间:
2011-03-01
影响因子:
3.7
通讯作者:
Wei, Zhi
Wei, Zhi
中科院分区:
数学1区
文献类型:
--
作者:
Sun, Wenguang;Wei, Zhi

文献摘要

被引文献

相似文献

在单课程实验中,通常需要识别出随时间推移而表现出特定差异表达模式的基因,从而深入了解潜在生物过程的机制。模式识别问题中的两个具有挑战性的问题是:(i)如何组合跨多个时间点的同时推断;(ii)如何在考虑强依赖性的同时控制多重性。我们提出了一个集智能多重测试的复合决策理论框架,并提出了一个数据驱动的过程,该过程的目的是在对假集率的约束下最小化遗漏集率。Yuan和Kendziorski(2006)提出的隐马尔可夫模型被推广到捕捉基因表达数据中的时间相关性。理论和数值结果表明,我们的数据驱动程序控制了多重性,提供了一种跨多个时间点组合同时推理的最佳方法,大大改进了传统的组合p值方法。特别地,我们展示了我们的方法在人类全身性炎症研究中的应用,用于检测早期和晚期反应基因。
In One-course experiments, it is often desirable to identify genes that exhibit a specific pattern of differential expression over time and thus gain insights into the mechanisms of the underlying biological processes. Two challenging issues in the pattern identification problem are: (i) how to combine the simultaneous inferences across multiple time points and (ii) how to control the multiplicity while accounting for the strong dependence. We formulate a compound decision-theoretic framework for set-wise multiple testing and propose a data-driven procedure that aims to minimize the missed set rate subject to a constraint on the false set rate. The hidden Markov model proposed in Yuan and Kendziorski (2006) is generalized to capture the temporal correlation in the gene expression data. Both theoretical and numerical results are presented to show that our data-driven procedure controls the multiplicity, provides an optimal way of combining simultaneous inferences across multiple time points, and greatly improves the conventional combined p-value methods. In particular, we demonstrate our method in an application to a study of systemic inflammation in humans for detecting early and late response genes.