Estimation of a significance threshold for genome-wide association studies

Estimation of a significance threshold for genome-wide association studies
复制标题

DOI:
10.1186/s12864-019-5992-7
复制
发表时间:
2019-07-29
期刊:
影响因子:
4.4
通讯作者:
Purcell, Larry C.
Purcell, Larry C.
中科院分区:
生物学2区
文献类型:
--
作者:
Kaler, Avjinder S.;Purcell, Larry C.

文献摘要

被引文献

相似文献

背景:在全基因组关联研究中,选择合适的统计显著性阈值对于区分真阳性、假阳性和假阴性至关重要。不同的多重检验比较方法已被开发来确定显著性阈值;然而,这些方法可能过于保守,可能导致假阴性增加。在这里,我们开发了一个经验公式,以确定统计显著性阈值,这是基于标记为基础的遗传力的性状。为了建立显著性阈值公式,我们在大豆、玉米和水稻中使用了45个模拟性状,这些性状在广义遗传力和qtl数量上都存在差异。结果采用1个自变量(基于标记的遗传力)和1个响应变量(- log(10) (P))的回归方程,建立显著性阈值的计算公式。对于所有物种,阈值-log(10) (P)值随着标记遗传力和广义遗传力的增加而增加。这些作物的广义遗传力较高,导致显著阈值较高。在作物种类中,玉米的连锁不平衡模式较低,其显著阈值高于大豆和水稻。结论与错误发现率和Bonferroni校正法相比,sour公式保守性更低,识别出更多的真阳性关联。
BackgroundSelection of an appropriate statistical significance threshold in genome-wide association studies is critical to differentiate true positives from false positives and false negatives. Different multiple testing comparison methods have been developed to determine the significance threshold; however, these methods may be overly conservative and may lead to an increase in false negatives. Here, we developed an empirical formula to determine the statistical significance threshold that is based on the marker-based heritability of the trait. To develop a formula for a significance threshold, we used 45 simulated traits in soybean, maize, and rice that varied in both broad sense heritability and the number of QTLs.ResultsA formula to determine a significance threshold was developed based on a regression equation that used one independent variable, marker-based heritability, and one response variable, - log(10) (P)-values. For all species, the threshold -log(10) (P)-values increased as both marker-based and broad-sense heritability increased. Higher broad sense heritability in these crops resulted in higher significant threshold values. Among crop species, maize, with a lower linkage disequilibrium pattern, had higher significant threshold values as compared to soybean and rice.ConclusionsOur formula was less conservative and identified more true positive associations than the false discovery rate and Bonferroni correction methods.