A database for efficient storage and management of multi panel SNP data

A database for efficient storage and management of multi panel SNP data
复制标题

DOI:
10.7482/0003-9438-56-103
复制
发表时间:
2013-11-20
影响因子:
--
通讯作者:
Truong, Cong V. C.
Truong, Cong V. C.
中科院分区:
农林科学3区
文献类型:
--
作者:
Groeneveld, Eildert;Truong, Cong V. C.

文献摘要

被引文献

相似文献

高通量基因分型的快速发展为遗传学开辟了新的可能性,同时也产生了巨大的数据处理问题。提出了一种系统设计和概念验证实现,它在关系数据库中提供单核苷酸多态性(SNP)基因型的高效数据存储和操作。使用SNP和个体选择向量的新策略允许我们将SNP数据视为矩阵或集合。这些基因型集提供了一种简单的方法来处理原始和衍生数据,后者基本上没有存储成本。由于其基于矢量的数据库存储,数据导入和导出比其他SNP数据库快得多。在概念验证实施中,压缩存储方案将磁盘空间需求减少了约300倍。此外,该设计与所涉及的个体和SNP的数量线性缩放。该过程支持不同大小的面板。这允许在同一种群中直接管理不同的面板大小,因为当较高密度的面板取代先前的较低密度版本时,它发生在动物育种计划中。
The fast development of high throughput genotyping has opened up new possibilities in genetics while at the same time producing immense data handling issues. A system design and proof of concept implementation are presented which provides efficient data storage and manipulation of single nucleotide polymorphism (SNP) genotypes in a relational database. A new strategy using SNP and individual selection vectors allows us to view SNP data as matrices or sets. These genotype sets provide an easy way to handle original and derived data, the latter at basically no storage costs. Due to its vector based database storage, data imports and exports are much faster than those of other SNP databases. In the proof of concept implementation, the compressed storage scheme reduces disk space requirements by a factor of around 300. Furthermore, this design scales linearly with number of individuals and SNPs involved. The procedure supports panels of different sizes. This allows a straight forward management of different panel sizes in the same population as it occurs in animal breeding programs when higher density panels replace previous lower density versions.