A modified generalized Fisher method for combining probabilities from dependent tests.

A modified generalized Fisher method for combining probabilities from dependent tests.
复制标题

DOI:
10.3389/fgene.2014.00032
复制
发表时间:
2014
影响因子:
3.7
通讯作者:
Cui Y
Cui Y
中科院分区:
生物学3区
文献类型:
--
作者:
Dai H;Leeder JS;Cui Y

文献摘要

参考文献

被引文献

相似文献

分子技术的快速发展已经产生了大量的高通量遗传数据来理解复杂性状的机理。基因变异的增加要求在分析中同时进行成百上千的统计测试,这对控制总体类型I错误率提出了挑战。组合来自多个假设检验的p值已显示出在高维遗传数据分析中聚集效应的前景。已经开发了几种p值组合方法并将其应用于遗传数据;见Dai等人。进行全面审查。然而,缺乏对相关遗传数据的调查,尤其是加权p值组合方法。由于连锁不平衡(LD),单核苷酸多态(SNPs)常常是相关的。其他遗传数据,包括来自下一代测序的变体、通过微阵列测量的基因表达水平、蛋白质和DNA甲基化数据等也包含复杂的关联结构。忽略遗传变量之间的相关结构可能会导致p值综合测试的I型错误率严重膨胀。在这项工作中,我们通过考虑p值之间的相关性结构,对Lancaster过程提出了改进。Lancaster过程中的权重函数允许将有意义的生物信息纳入统计分析,这可以增加统计测试的能力和/或消除过程中的偏差。广泛的经验评估表明,改进的Lancaster过程大大降低了由于p值之间的相关性而导致的I类错误率,并保持了相当大的检测p值之间的信号的能力。我们应用我们的方法重新评估了已发表的肾移植数据,并确定了B细胞通路和同种异体移植耐受之间的新关联。
Rapid developments in molecular technology have yielded a large amount of high throughput genetic data to understand the mechanism for complex traits. The increase of genetic variants requires hundreds and thousands of statistical tests to be performed simultaneously in analysis, which poses a challenge to control the overall Type I error rate. Combining p-values from multiple hypothesis testing has shown promise for aggregating effects in high-dimensional genetic data analysis. Several p-value combining methods have been developed and applied to genetic data; see Dai et al. for a comprehensive review. However, there is a lack of investigations conducted for dependent genetic data, especially for weighted p-value combining methods. Single nucleotide polymorphisms (SNPs) are often correlated due to linkage disequilibrium (LD). Other genetic data, including variants from next generation sequencing, gene expression levels measured by microarray, protein and DNA methylation data, etc. also contain complex correlation structures. Ignoring correlation structures among genetic variants may lead to severe inflation of Type I error rates for omnibus testing of p-values. In this work, we propose modifications to the Lancaster procedure by taking the correlation structure among p-values into account. The weight function in the Lancaster procedure allows meaningful biological information to be incorporated into the statistical analysis, which can increase the power of the statistical testing and/or remove the bias in the process. Extensive empirical assessments demonstrate that the modified Lancaster procedure largely reduces the Type I error rates due to correlation among p-values, and retains considerable power to detect signals among p-values. We applied our method to reassess published renal transplant data, and identified a novel association between B cell pathways and allograft tolerance.
DOI: 10.1086/522374
发表时间: 2007-12-01
影响因子: 9.8
作者:
Wang, Kai;Li, Mingyao;Bucan, Maja
通讯作者: Bucan, Maja
DOI: 10.2307/2340521
发表时间: 1922-01-01
影响因子: --
作者:
Fisher, RA
通讯作者: Fisher, RA
DOI: 10.1214/aoms/1177698949
发表时间: 1967-01-01
影响因子: --
作者:
BAHADUR, RR
通讯作者: BAHADUR, RR
DOI: 10.1080/02664760701683528
发表时间: 2008-01-01
影响因子: 1.5
作者:
Dai, Hongying;Charnigo, Richard
通讯作者: Charnigo, Richard
DOI: 10.1002/bimj.4710380603
发表时间: 1996-01-01
影响因子: 1.7
作者:
Koziol, JA
通讯作者: Koziol, JA