Scalable probabilistic PCA for large-scale genetic variation data

Scalable probabilistic PCA for large-scale genetic variation data
复制标题

DOI:
10.1371/journal.pgen.1008773
复制
发表时间:
2020-05-01
期刊:
影响因子:
4.5
通讯作者:
Sankararaman, Sriram
Sankararaman, Sriram
中科院分区:
生物学2区
文献类型:
--
作者:
Agrawal, Aman;Chiu, Alec M.;Sankararaman, Sriram

文献摘要

被引文献

相似文献

主成分分析是了解群体结构和遗传变异的常用技术。随着包含数十万个体遗传信息的大规模数据集的出现,需要能够计算具有可扩展计算和存储要求的主成分(PC)的方法。在这项研究中,我们提出了ProPCA,一个高度可扩展的统计方法来有效地计算遗传PC。我们系统地评估了我们的方法在大规模模拟数据上的准确性和可扩展性,并将其应用于英国生物银行。利用ProPCA推断的英国生物库中的白色英国个体的群体结构,我们确定了几个新的信号,推定最近selection.PCA(主成分分析)是一个关键的工具,了解人口结构和控制人口分层的全基因组关联研究(GWAS)。随着大规模遗传变异数据集的出现,需要能够计算具有可扩展计算和存储要求的主成分(PC)的方法。我们提出了ProPCA,一个高度可扩展的方法的基础上的概率生成模型,有效地计算遗传变异数据的前PC。我们应用ProPCA在大约30分钟内计算了来自英国生物银行的基因型数据的前五个PC,包括488,363个个体和146,671个SNP。为了说明在大样本中计算PC的实用性,我们利用英国生物库中白色英国个体中ProPCA推断的群体结构来确定最近推定选择的几个新的全基因组信号,包括RPGRIP1L和TLR 4中的错义突变。
Author summaryPrincipal component analysis is a commonly used technique for understanding population structure and genetic variation. With the advent of large-scale datasets that contain the genetic information of hundreds of thousands of individuals, there is a need for methods that can compute principal components (PCs) with scalable computational and memory requirements. In this study, we present ProPCA, a highly scalable statistical method to compute genetic PCs efficiently. We systematically evaluate the accuracy and scalability of our method on large-scale simulated data and apply it to the UK Biobank. Leveraging the population structure inferred by ProPCA within the White British individuals in the UK Biobank, we identify several novel signals of putative recent selection.Principal component analysis (PCA) is a key tool for understanding population structure and controlling for population stratification in genome-wide association studies (GWAS). With the advent of large-scale datasets of genetic variation, there is a need for methods that can compute principal components (PCs) with scalable computational and memory requirements. We present ProPCA, a highly scalable method based on a probabilistic generative model, which computes the top PCs on genetic variation data efficiently. We applied ProPCA to compute the top five PCs on genotype data from the UK Biobank, consisting of 488,363 individuals and 146,671 SNPs, in about thirty minutes. To illustrate the utility of computing PCs in large samples, we leveraged the population structure inferred by ProPCA within White British individuals in the UK Biobank to identify several novel genome-wide signals of recent putative selection including missense mutations in RPGRIP1L and TLR4.