Efficient analysis of large-scale genome-wide data with two R packages: bigstatsr and bigsnpr.

Efficient analysis of large-scale genome-wide data with two R packages: bigstatsr and bigsnpr.
复制标题

DOI:
10.1093/bioinformatics/bty185
复制
发表时间:
2018-08-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Blum MGB
Blum MGB
中科院分区:
其他
文献类型:
--
作者:
Privé F;Aschard H;Ziyatdinov A;Blum MGB

文献摘要

参考文献

被引文献

相似文献

在过去几年中,为关联研究生成的全基因组数据集的规模急剧增加,现代数据集通常包括在数十万人中测量的数百万个变异。数据量的增加是一个重大挑战,严重减慢了基因组分析的速度,导致一些软件变得过时,并且研究人员对各种分析工具的访问受到限制。在这里,我们展示了两个 R 包:bigstatsr 和 bigsnpr,允许在 R 中执行大规模基因组数据的分析。为了解决大数据量问题,这些包使用内存映射来访问存储在磁盘而不是 RAM 中的数据矩阵。为了执行数据预处理和数据分析,这些软件包集成了大多数常用工具,或者通过对现有软件的透明系统调用,或者通过更新或改进现有方法的实现。特别是,这些软件包实现了主成分分析和关联研究的快速准确计算,消除连锁不平衡中的单核苷酸多态性的功能以及学习数百万个单核苷酸多态性的多基因风险评分的算法。我们通过分析乳糜泻病例对照基因组数据集、进行关联研究和计算多基因风险评分来说明这两个 R 包的应用。最后,我们通过分析模拟的全基因组数据集(包括一台台式计算机上的 500 000 个人和 100 万个标记)来展示 R 包的可扩展性。 https://privefl.github.io/bigstatsr/ 和 https://privefl.github.io/bigsnpr/。 补充数据可在生物信息学在线获取。
Genome-wide datasets produced for association studies have dramatically increased in size over the past few years, with modern datasets commonly including millions of variants measured in dozens of thousands of individuals. This increase in data size is a major challenge severely slowing down genomic analyses, leading to some software becoming obsolete and researchers having limited access to diverse analysis tools. Here we present two R packages, bigstatsr and bigsnpr, allowing for the analysis of large scale genomic data to be performed within R. To address large data size, the packages use memory-mapping for accessing data matrices stored on disk instead of in RAM. To perform data pre-processing and data analysis, the packages integrate most of the tools that are commonly used, either through transparent system calls to existing software, or through updated or improved implementation of existing methods. In particular, the packages implement fast and accurate computations of principal component analysis and association studies, functions to remove single nucleotide polymorphisms in linkage disequilibrium and algorithms to learn polygenic risk scores on millions of single nucleotide polymorphisms. We illustrate applications of the two R packages by analyzing a case–control genomic dataset for celiac disease, performing an association study and computing polygenic risk scores. Finally, we demonstrate the scalability of the R packages by analyzing a simulated genome-wide dataset including 500 000 individuals and 1 million markers on a single desktop computer. https://privefl.github.io/bigstatsr/ and https://privefl.github.io/bigsnpr/. Supplementary data are available at Bioinformatics online.
DOI: 10.1186/1471-2105-13-88
发表时间: 2012-05-10
期刊: BMC bioinformatics
影响因子: 3
作者:
Abraham G;Kowalczyk A;Zobel J;Inouye M
通讯作者: Inouye M
DOI: 10.1093/bioinformatics/btu848
发表时间: 2015-05-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Euesden J;Lewis CM;O'Reilly PF
通讯作者: O'Reilly PF
DOI: 10.1186/s13742-015-0047-8
发表时间: 2015
期刊: GigaScience
影响因子: 9.2
作者:
Chang CC;Chow CC;Tellier LC;Vattikuti S;Purcell SM;Lee JJ
通讯作者: Lee JJ
DOI: 10.1186/1471-2105-9-526
发表时间: 2008-12-08
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Nielsen, Jesper;Mailund, Thomas
通讯作者: Mailund, Thomas
DOI: 10.1186/1471-2105-14-166
发表时间: 2013-05-28
期刊: BMC bioinformatics
影响因子: 3
作者:
Sikorska K;Lesaffre E;Groenen PF;Eilers PH
通讯作者: Eilers PH