SDhaP: haplotype assembly for diploids and polyploids via semi-definite programming.

SDhaP: haplotype assembly for diploids and polyploids via semi-definite programming.
复制标题

DOI:
10.1186/s12864-015-1408-5
复制
发表时间:
2015-04-03
期刊:
影响因子:
4.4
通讯作者:
Vikalo H
Vikalo H
中科院分区:
生物学2区
文献类型:
--
作者:
Das S;Vikalo H

文献摘要

参考文献

被引文献

相似文献

单倍型组装的目标是从测序的染色体片段的混合物中推断个体的单倍型。有限长度的双端测序读段和插入片段使得单倍型组装在计算上具有挑战性;事实上,已知大多数问题公式是NP难的。随着测序技术的进步以及读段和插入片段长度的增长,单倍型组装问题的规模(以及因此的难度)不断增加。在多倍体单倍型的情况下,计算挑战甚至更加明显,其组装比二倍体的情况困难得多。需要用于二倍体和多倍体生物的单倍型组装的快速、准确和可扩展的方法。我们从高通量测序数据中开发了一种用于二倍体/多倍体单倍型组装的新框架。该方法将单倍型组装问题表述为半定规划,并利用其特殊结构-即底层解的低秩-来快速且高精度地解决该问题。开发的框架是适用于二倍体和多倍体物种。SDhaP的代码可以在https://sourceforge.net/projects/sdhap上免费获得。对真实的数据和模拟数据的大量基准测试表明,所提出的算法在准确性或速度或两者方面优于几种知名的单倍型组装方法。还提供了实现接近最优解决方案所需的覆盖率的有用建议。
The goal of haplotype assembly is to infer haplotypes of an individual from a mixture of sequenced chromosome fragments. Limited lengths of paired-end sequencing reads and inserts render haplotype assembly computationally challenging; in fact, most of the problem formulations are known to be NP-hard. Dimensions (and, therefore, difficulty) of the haplotype assembly problems keep increasing as the sequencing technology advances and the length of reads and inserts grow. The computational challenges are even more pronounced in the case of polyploid haplotypes, whose assembly is considerably more difficult than in the case of diploids. Fast, accurate, and scalable methods for haplotype assembly of diploid and polyploid organisms are needed. We develop a novel framework for diploid/polyploid haplotype assembly from high-throughput sequencing data. The method formulates the haplotype assembly problem as a semi-definite program and exploits its special structure – namely, the low rank of the underlying solution – to solve it rapidly and with high accuracy. The developed framework is applicable to both diploid and polyploid species. The code for SDhaP is freely available at https://sourceforge.net/projects/sdhap. Extensive benchmarking tests on both real and simulated data show that the proposed algorithms outperform several well-known haplotype assembly methods in terms of either accuracy or speed or both. Useful recommendations for coverages needed to achieve near-optimal solutions are also provided.
DOI: 10.1089/cmb.2012.0084
发表时间: 2012-06-01
影响因子: 1.7
作者:
Aguiar, Derek;Istrail, Sorin
通讯作者: Istrail, Sorin
使用下一代 DNA 测序数据进行变异发现和基因分型的框架。
DOI: 10.1038/ng.806
发表时间: 2011-05
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1093/bioinformatics/btq215
发表时间: 2010-06-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
He D;Choi A;Pipatsrisawat K;Darwiche A;Eskin E
通讯作者: Eskin E
DOI: 10.1038/nature02168
发表时间: 2003-12-18
期刊: NATURE
影响因子: 64.8
作者:
Gibbs, RA;Belmont, JW;Tanaka, T
通讯作者: Tanaka, T
DOI: 10.1101/gr.077065.108
发表时间: 2008-08-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Bansal, Vikas;Halpern, Aaron L.;Bafna, Vineet
通讯作者: Bafna, Vineet