Effect of non-normality and low count variants on cross-phenotype association tests in GWAS.

Effect of non-normality and low count variants on cross-phenotype association tests in GWAS.
复制标题

非正态性和低计数变异对 GWAS 交叉表型关联测试的影响。

DOI:
10.1038/s41431-019-0514-2
复制
发表时间:
2020
期刊:
European journal of human genetics : EJHG
影响因子:
--
通讯作者:
Chatterjee,Nilanjan
Chatterjee,Nilanjan
中科院分区:
--
文献类型:
--
作者:
Ray,Debashree;Chatterjee,Nilanjan

文献摘要

相似文献

许多复杂的人类疾病,如2型糖尿病,其特征是具有基本相同的遗传结构的多种潜在特征/表型。相关性状的多变量分析有可能增加检测潜在共同遗传位点的能力。已经提出了几种交叉表型关联方法--一些需要关于性状和基因类型的个体水平的数据,而另一些只需要总结水平的数据。在这篇文章中,我们探讨了多变量性状分布的非正态分布是否会影响现有的一些多性状方法的推断,以及这种影响是如何依赖于被测试遗传变量的等位基因数量的。我们发现,这些测试中的大多数都容易受到导致虚假关联信号的偏差的影响。即使在控制了可能导致非正态的混杂因素,然后对每个性状的残差进行反向正态转换后,这些测试可能会对次要等位基因计数(MAC)较低的变种产生夸大的I型错误。基于性状的个体水平基因型的有序回归的关联似然比检验似乎是最无偏见的,并且当MAC值相当大时(例如,MAC&>; 30)可以维持类型I错误。将这些方法应用于欧洲样本上可公开获得的8个氨基酸性状的汇总统计数据,似乎显示出系统的膨胀(特别是对于MAC较低的变种),这与我们从模拟实验中的发现一致。
Many complex human diseases, such as type 2 diabetes, are characterized by multiple underlying traits/phenotypes that have substantially shared genetic architecture. Multivariate analysis of correlated traits has the potential to increase the power of detecting underlying common genetic loci. Several cross-phenotype association methods have been proposed—some require individual-level data on traits and genotypes, while the others require only summary-level data. In this article, we explore whether non-normality of multivariate trait distribution affects the inference from some of the existing multi-trait methods and how that effect is dependent on the allele count of the genetic variant being tested. We find that most of these tests are susceptible to biases that lead to spurious association signals. Even after controlling for confounders that may contribute to non-normality and then applying inverse normal transformation on the residuals of each trait, these tests may have inflated type I errors for variants with low minor allele counts (MACs). A likelihood ratio test of association based on the ordinal regression of individual-level genotype conditional on the traits seems to be the least biased and can maintain type I error when the MAC is reasonably large (e.g., MAC > 30). Application of these methods to publicly available summary statistics of eight amino acid traits on European samples seem to exhibit systematic inflation (especially for variants with low MAC), which is consistent with our findings from simulation experiments.