FiMAP: A fast identity-by-descent mapping test for biobank-scale cohorts.

FiMAP: A fast identity-by-descent mapping test for biobank-scale cohorts.
复制标题

DOI:
10.1371/journal.pgen.1011057
复制
发表时间:
2023-12
期刊:
影响因子:
4.5
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

虽然全基因组关联研究(GWAS)已经确定了数以万计的遗传位点,但许多复杂性状的遗传结构仍然没有完全理解。大多数GWAS和测序关联研究都集中在单核苷酸多态性或拷贝数变异上,包括常见和罕见的遗传变异。然而,在GWAS或罕见变异的变异集测试中,定相单倍型信息经常被忽略。在这里,我们利用从基于随机投影的IBD检测算法中推断出的后代身份(IBD)片段,在具有复杂性状的遗传关联的映射中,为生物库规模的队列中的IBD映射开发一种计算效率高的统计测试。我们使用稀疏线性代数和随机矩阵算法来加快计算速度,在几个小时内完成了超过40万个样本的全基因组IBD映射扫描。模拟研究表明,我们的新方法在大型生物库规模队列中没有遗传关联的零假设下具有良好的控制I型错误率,并且在因果变异未分型和罕见或存在单倍型效应时优于传统的GWAS单变异检验。我们还将我们的方法应用于使用英国生物银行数据的六个人体测量特征的IBD映射,并确定了总共3,442个关联,其中2,131个(62%)在GWAS的± 3厘摩侧翼区域中的暗示性标签变体条件化后仍然显着。从共同祖先共享的血统同一性(IBD)片段可用于关联作图以鉴定与复杂性状(如身高、体重和体脂分布)相关的基因组区域。在这项工作中,我们开发了FiMAP,这是一种有效的IBD映射测试,可以通过利用最近有效的IBD检测软件程序中的IBD片段,应用于来自生物库规模队列的数十万个体。FiMAP建立在经典的方差分量模型的基础上,该模型在连锁分析中有很深的根源,FiMAP利用稀疏线性代数和随机矩阵算法来加快大样本的计算速度。我们证明了FiMAP算法的准确性,I型错误控制和统计功效,并说明了使用FiMAP的IBD映射如何识别与全基因组关联研究结果互补的关联。随着大型生物库规模队列中丰富的遗传数据的可用性,FiMAP提供了一个独特的角度来利用这些丰富的资源,更好地了解复杂性状的遗传结构。
Although genome-wide association studies (GWAS) have identified tens of thousands of genetic loci, the genetic architecture is still not fully understood for many complex traits. Most GWAS and sequencing association studies have focused on single nucleotide polymorphisms or copy number variations, including common and rare genetic variants. However, phased haplotype information is often ignored in GWAS or variant set tests for rare variants. Here we leverage the identity-by-descent (IBD) segments inferred from a random projection-based IBD detection algorithm in the mapping of genetic associations with complex traits, to develop a computationally efficient statistical test for IBD mapping in biobank-scale cohorts. We used sparse linear algebra and random matrix algorithms to speed up the computation, and a genome-wide IBD mapping scan of more than 400,000 samples finished within a few hours. Simulation studies showed that our new method had well-controlled type I error rates under the null hypothesis of no genetic association in large biobank-scale cohorts, and outperformed traditional GWAS single-variant tests when the causal variants were untyped and rare, or in the presence of haplotype effects. We also applied our method to IBD mapping of six anthropometric traits using the UK Biobank data and identified a total of 3,442 associations, 2,131 (62%) of which remained significant after conditioning on suggestive tag variants in the ± 3 centimorgan flanking regions from GWAS. Identity-by-descent (IBD) segments shared from a common ancestor can be used in association mapping to identify genomic regions associated with complex traits such as height, weight and body fat distribution. In this work, we developed FiMAP, an efficient IBD mapping test that can be applied to hundreds of thousands of individuals from biobank-scale cohorts, by leveraging IBD segments from recent efficient IBD detection software programs. Built upon the classical variance component model which has its deep root in linkage analysis, FiMAP utilizes sparse linear algebra and random matrix algorithms to speed up the computation in large samples. We demonstrated accuracy, type I error control and statistical power of the FiMAP algorithm, and illustrated how IBD mapping using FiMAP could identify associations complementary to genome-wide association study findings. With the availability of an abundance of genetic data from large biobank-scale cohorts, FiMAP provides a unique angle to leverage such rich resources and better understand the genetic architecture of complex traits.
DOI: 10.1038/ejhg.2014.155
发表时间: 2015-05
期刊: European journal of human genetics : EJHG
影响因子: --
作者:
通讯作者: --
DOI: 10.1002/0471142905.hg0128s84
发表时间: 2015-01-20
影响因子: --
作者:
Thornton, Timothy A
通讯作者: Thornton, Timothy A
DOI: 10.1186/s12915-021-00964-y
发表时间: 2021-02-16
期刊: BMC biology
影响因子: 5.4
作者:
Naseri A;Tang K;Geng X;Shi J;Zhang J;Shakya P;Liu X;Zhang S;Zhi D
通讯作者: Zhi D