Imputation of Truncated p-Values For Meta-Analysis Methods and Its Genomic Application.

Imputation of Truncated p-Values For Meta-Analysis Methods and Its Genomic Application.
复制标题

DOI:
10.1214/14-aoas747
复制
发表时间:
2014-12
期刊:
The annals of applied statistics
影响因子:
--
通讯作者:
Tseng GC
Tseng GC
中科院分区:
其他
文献类型:
--
作者:
Tang S;Ding Y;Sibille E;Mogil J;Lariviere WR;Tseng GC

文献摘要

相似文献

在过去的十年中,同时监测数千个基因表达活动的微阵列分析已成为生物医学研究的常规方法。大量的表达谱被生成并存储在公共领域,并且通过荟萃分析来检测差异表达(DE)基因的信息整合已变得流行,以获得增强的统计能力和经过验证的结果。聚合转化 p 值证据的方法已广泛应用于基因组环境中,其中 Fisher 方法和 Stouffer 方法是最流行的方法。在实践中,DE 证据的原始数据和 p 值通常无法在需要合并的基因组研究中获得。相反,期刊出版物中仅报告在特定 p 值阈值下检测到的 DE 基因列表(例如,p 值 < 0.001 的 DE 基因)。截断的p值信息使得上述荟萃分析方法不再适用,研究人员被迫采用效率较低的计票方法,或者天真地放弃信息不完整的研究。本文的目的是针对这种 p 值部分删失的情况开发有效的荟萃分析方法。我们针对一般类别的证据聚合方法开发并比较了三种插补方法——平均插补、单次随机插补和多重插补,其中费舍尔和斯托弗的方法是特殊的例子。分析得出每种方法的零分布,并建立随后的推断和基因组分析框架。进行模拟以研究(相关)基因表达数据的 I 类错误、功效和错误发现率 (FDR) 的控制。所提出的方法已应用于结直肠癌、疼痛和重度抑郁症(MDD)的液体关联分析中的多种基因组应用。结果表明,插补方法优于现有的朴素方法。平均插补和多重插补方法表现最好,推荐用于未来的应用。
Microarray analysis to monitor expression activities in thousands of genes simultaneously has become routine in biomedical research during the past decade. a tremendous amount of expression profiles are generated and stored in the public domain and information integration by meta-analysis to detect differentially expressed (DE) genes has become popular to obtain increased statistical power and validated findings. Methods that aggregate transformed p-value evidence have been widely used in genomic settings, among which Fisher's and Stouffer's methods are the most popular ones. In practice, raw data and p-values of DE evidence are often not available in genomic studies that are to be combined. Instead, only the detected DE gene lists under a certain p-value threshold (e.g., DE genes with p-value < 0.001) are reported in journal publications. The truncated p-value information makes the aforementioned meta-analysis methods inapplicable and researchers are forced to apply a less efficient vote counting method or naïvely drop the studies with incomplete information. The purpose of this paper is to develop effective meta-analysis methods for such situations with partially censored p-values. We developed and compared three imputation methods—mean imputation, single random imputation and multiple imputation—for a general class of evidence aggregation methods of which Fisher's and Stouffer's methods are special examples. The null distribution of each method was analytically derived and subsequent inference and genomic analysis frameworks were established. Simulations were performed to investigate the type Ierror, power and the control of false discovery rate (FDR) for (correlated) gene expression data. The proposed methods were applied to several genomic applications in colorectal cancer, pain and liquid association analysis of major depressive disorder (MDD). The results showed that imputation methods outperformed existing naïve approaches. Mean imputation and multiple imputation methods performed the best and are recommended for future applications.