TeraPCA: a fast and scalable software package to study genetic variation in tera-scale genotypes
TeraPCA: a fast and scalable software package to study genetic variation in tera-scale genotypes
复制标题
DOI:
10.1093/bioinformatics/btz157
复制
发表时间:
2019-10-01
期刊:
影响因子:
5.8
通讯作者:
Drineas, Petros
中科院分区:
文献类型:
--
作者:
Bose, Aritra;Kalantzis, Vassilis;Drineas, Petros
Motivation: Principal Component Analysis is a key tool in the study of population structure in human genetics. As modern datasets become increasingly larger in size, traditional approaches based on loading the entire dataset in the system memory (Random Access Memory) become impractical and out-of-core implementations are the only viable alternative.Results: We present TeraPCA, a C++ implementation of the Randomized Subspace Iteration method to perform Principal Component Analysis of large-scale datasets. TeraPCA can be applied both in-core and out-of-core and is able to successfully operate even on commodity hardware with a system memory of just a few gigabytes. Moreover, TeraPCA has minimal dependencies on external libraries and only requires a working installation of the BLAS and LAPACK libraries. When applied to a dataset containing a million individuals genotyped on a million markers, TeraPCA requires