Efficient Variant Set Mixed Model Association Tests for Continuous and Binary Traits in Large-Scale Whole-Genome Sequencing Studies

Efficient Variant Set Mixed Model Association Tests for Continuous and Binary Traits in Large-Scale Whole-Genome Sequencing Studies
复制标题

DOI:
10.1016/j.ajhg.2018.12.012
复制
发表时间:
2019-02-07
影响因子:
9.8
通讯作者:
Lin, Xihong
Lin, Xihong
中科院分区:
生物学1区
文献类型:
--
作者:
Chen, Han;Huffman, Jennifer E.;Lin, Xihong

文献摘要

被引文献

相似文献

随着全基因组测序(WGS)技术的进步,正在开发更先进的统计方法来测试与罕见变异的遗传关联。将变体分组进行分析的方法也称为变体集、基于基因的和聚合单元测试。负担检验和序列核关联检验(SKAT)是两种广泛使用的变异集检验,最初是为无关个体的样本开发的,后来扩展到具有已知系谱结构的家系数据。然而,需要计算效率高和功能强大的变量集测试,使分析易于处理的大规模WGS研究与复杂的研究样本。在广义线性混合模型框架下,提出了连续性状和二元性状的变异集混合模型关联检验(SMMAT)。这些测试可以应用于涉及具有群体结构和相关性的样本的大规模WGS研究,例如国家心脏,肺和血液研究所的精准医学跨组学(TOPMed)计划。SMMAT对于不同的变异集共享相同的空模型,并且该空模型(其仅包括协变量)的优点在于,对于每个全基因组分析中的所有测试,它仅需要拟合一次。仿真研究表明,所有建议的SMMAT正确控制I型错误率的连续和二进制性状的存在下,人口结构和相关性。我们还说明了我们的测试在一个真实的数据的例子,分析血浆纤维蛋白原水平的TOPMed程序(n = 23,763),使用分析共享,基于云的计算平台。
With advances in whole-genome sequencing (WGS) technology, more advanced statistical methods for testing genetic association with rare variants are being developed. Methods in which variants are grouped for analysis are also known as variant-set, gene-based, and aggregate unit tests. The burden test and sequence kernel association test (SKAT) are two widely used variant-set tests, which were originally developed for samples of unrelated individuals and later have been extended to family data with known pedigree structures. However, computationally efficient and powerful variant-set tests are needed to make analyses tractable in large-scale WGS studies with complex study samples. In this paper, we propose the variant-set mixed model association tests (SMMAT) for continuous and binary traits using the generalized linear mixed model framework. These tests can be applied to large-scale WGS studies involving samples with population structure and relatedness, such as in the National Heart, Lung, and Blood Institute's Trans-Omics for Precision Medicine (TOPMed) program. SMMATs share the same null model for different variant sets, and a virtue of this null model, which includes covariates only, is that it needs to be fit only once for all tests in each genome-wide analysis. Simulation studies show that all the proposed SMMATs correctly control type I error rates for both continuous and binary traits in the presence of population structure and relatedness. We also illustrate our tests in a real data example of analysis of plasma fibrinogen levels in the TOPMed program (n = 23,763), using the Analysis Commons, a cloud-based computing platform.