TheSNPpit-A High Performance Database System for Managing Large Scale SNP Data.

TheSNPpit-A High Performance Database System for Managing Large Scale SNP Data.
复制标题

thesnppit-A高性能数据库系统用于管理大型SNP数据。

DOI:
10.1371/journal.pone.0164043
复制
发表时间:
2016
期刊:
影响因子:
3.7
通讯作者:
Lichtenberg H
Lichtenberg H
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Groeneveld E;Lichtenberg H

文献摘要

参考文献

被引文献

相似文献

高通量基因分型的快速发展为遗传学开辟了新的可能性,同时也产生了相当多的数据处理问题。TheSNPit是一个数据库系统,用于管理来自任何基因分型平台的大量多组SNP基因型数据。随着动植物育种以及人类遗传学等领域的基因分型率不断提高,现在已经有数十万人需要管理。虽然每个SNP一行的常见数据库设计可以管理数百个样本,但随着数据集大小的增加,这种方法变得越来越慢,直到最终在需要管理数万甚至数十万个个体时完全失败。TheSNPit实现了三个想法来适应这样的大规模实验:在关系数据库中高度压缩的向量存储,基于集合的数据操作,以及用C编写的非常快速的导出,Perl作为框架的基础,PostgreSQL作为数据库后端。其新颖的子集系统允许基于SNP过滤(基于主要等位基因频率,无调用和染色体)和手动应用的样本和SNP列表创建命名子集,存储成本可以忽略不计,从而避免了文件副本激增的问题。命名的子集被导出用于下游分析。PLINK ped和映射文件作为输入和输出进行处理。在动物和植物育种计划中,当较高密度的面板取代先前的较低密度版本时,SNPit允许在相同的个体群体中管理不同的面板大小。一个完全通用的程序允许存储表型。SNP pit仅占用2位用于存储单个SNP,这意味着每1 MB磁盘存储容量为4 mio SNP。为了研究性能扩展,已经创建了一个具有超过1850万个样本的数据库,该数据库具有来自12个组的3.4万亿个SNP,范围从1000到2000万个SNP,从而产生850 GB的数据库。导入和导出性能与SNP的数量呈线性关系,并且在很大程度上与面板和数据库大小无关。导入速度约为6 mio SNP/sec,导出速度在60到120 mio SNP/sec之间。基于命令行,导入和导出可以轻松集成到管道中。SNPit在开源GNU通用公共许可证(GPL)第2版下可用。
The fast development of high throughput genotyping has opened up new possibilities in genetics while at the same time producing considerable data handling issues. TheSNPpit is a database system for managing large amounts of multi panel SNP genotype data from any genotyping platform. With an increasing rate of genotyping in areas like animal and plant breeding as well as human genetics, already now hundreds of thousand of individuals need to be managed. While the common database design with one row per SNP can manage hundreds of samples this approach becomes progressively slower as the size of the data sets increase until it finally fails completely once tens or even hundreds of thousands of individuals need to be managed. TheSNPpit has implemented three ideas to also accomodate such large scale experiments: highly compressed vector storage in a relational database, set based data manipulation, and a very fast export written in C with Perl as the base for the framework and PostgreSQL as the database backend. Its novel subset system allows the creation of named subsets based on the filtering of SNP (based on major allele frequency, no-calls, and chromosomes) and manually applied sample and SNP lists at negligible storage costs, thus avoiding the issue of proliferating file copies. The named subsets are exported for down stream analysis. PLINK ped and map files are processed as in- and outputs. TheSNPpit allows management of different panel sizes in the same population of individuals when higher density panels replace previous lower density versions as it occurs in animal and plant breeding programs. A completely generalized procedure allows storage of phenotypes. TheSNPpit only occupies 2 bits for storing a single SNP implying a capacity of 4 mio SNPs per 1MB of disk storage. To investigate performance scaling, a database with more than 18.5 mio samples has been created with 3.4 trillion SNPs from 12 panels ranging from 1000 through 20 mio SNPs resulting in a database of 850GB. The import and export performance scales linearly with the number of SNPs and is largely independent of panel and database size. Import speed is around 6 mio SNPs/sec, export between 60 and 120 mio SNPs/sec. Being command line based, imports and exports can easily be integrated into pipelines. TheSNPpit is available under the Open Source GNU General Public License (GPL) Version 2.
DOI: 10.1093/database/bau098
发表时间: 2014-10-03
影响因子: 5.8
作者:
Ameur, Adam;Bunikis, Ignas;Gyllensten, Ulf
通讯作者: Gyllensten, Ulf
DOI: 10.1093/bioinformatics/btp714
发表时间: 2010-02-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Fong C;Ko DC;Wasnick M;Radey M;Miller SI;Brittnacher M
通讯作者: Brittnacher M
DOI: 10.7482/0003-9438-56-103
发表时间: 2013-11-20
影响因子: --
作者:
Groeneveld, Eildert;Truong, Cong V. C.
通讯作者: Truong, Cong V. C.
DOI: 10.1186/1471-2105-11-238
发表时间: 2010-05-11
期刊: BMC bioinformatics
影响因子: 3
作者:
Rios D;McLaren WM;Chen Y;Birney E;Stabenau A;Flicek P;Cunningham F
通讯作者: Cunningham F
DOI: 10.1007/978-1-62703-447-0_4
发表时间: 2013-01-01
期刊: Methods in molecular biology (Clifton, N.J.)
影响因子: --
作者:
Mitha, Faheem
通讯作者: Mitha, Faheem