TeraPCA: a fast and scalable software package to study genetic variation in tera-scale genotypes

TeraPCA: a fast and scalable software package to study genetic variation in tera-scale genotypes
复制标题

DOI:
10.1093/bioinformatics/btz157
复制
发表时间:
2019-10-01
期刊:
影响因子:
5.8
通讯作者:
Drineas, Petros
Drineas, Petros
中科院分区:
生物学3区
文献类型:
--
作者:
Bose, Aritra;Kalantzis, Vassilis;Drineas, Petros

文献摘要

被引文献

相似文献

动机:主成分分析是研究人类遗传学群体结构的关键工具。随着现代数据集的规模越来越大,基于将整个数据集加载到系统内存(随机访问内存)的传统方法变得不切实际,而核外实现是唯一可行的替代方案。结果:我们提出了TeraPCA,一个c++实现的随机子空间迭代方法,用于执行大规模数据集的主成分分析。TeraPCA既可以应用于核内也可以应用于核外,即使在系统内存只有几gb的商用硬件上也能成功运行。此外,TeraPCA对外部库的依赖最小,只需要BLAS和LAPACK库的工作安装。当应用于包含一百万个个体基因分型在一百万个标记的数据集时,TeraPCA需要
Motivation: Principal Component Analysis is a key tool in the study of population structure in human genetics. As modern datasets become increasingly larger in size, traditional approaches based on loading the entire dataset in the system memory (Random Access Memory) become impractical and out-of-core implementations are the only viable alternative.Results: We present TeraPCA, a C++ implementation of the Randomized Subspace Iteration method to perform Principal Component Analysis of large-scale datasets. TeraPCA can be applied both in-core and out-of-core and is able to successfully operate even on commodity hardware with a system memory of just a few gigabytes. Moreover, TeraPCA has minimal dependencies on external libraries and only requires a working installation of the BLAS and LAPACK libraries. When applied to a dataset containing a million individuals genotyped on a million markers, TeraPCA requires