WHATSHAP: Weighted Haplotype Assembly for Future-Generation Sequencing Reads

WHATSHAP: Weighted Haplotype Assembly for Future-Generation Sequencing Reads
复制标题

DOI:
10.1089/cmb.2014.0157
复制
发表时间:
2015-06-01
影响因子:
1.7
通讯作者:
Schonhuth, Alexander
Schonhuth, Alexander
中科院分区:
生物学4区
文献类型:
--
作者:
Patterson, Murray;Marschall, Tobias;Schonhuth, Alexander

文献摘要

被引文献

相似文献

人类基因组是二倍体,这需要将杂合单核苷酸多态性(SNP)分配给基因组的两个拷贝。由此产生的单倍型,属于每个拷贝的SNPs列表,对于群体遗传学的下游分析至关重要。目前,统计学方法,这是无视直接读取信息,构成了国家的最先进的。单倍型组装,它解决了直接从测序读数定相,遭受的事实,即目前一代的测序读数太短,服务于全基因组定相的目的。虽然未来技术测序读数将包含足够量的SNP/读数用于定相,但它们也可能遭受更高的测序错误率。目前,不存在允许同时考虑增加的读段长度和测序错误信息的单体型组装方法。在这里,我们建议WhatsHap,第一种方法,产生可证明的最佳解决方案的加权最小纠错问题的运行时线性的SNP的数量。WhatsHap是一种固定参数易处理(FPT)方法,以覆盖率为参数。我们证明了WhatsHap可以处理覆盖率高达20倍的数据集,并且即使在测序错误率显著升高的情况下,15倍通常也足以可靠地对长读段进行定相。我们还发现,我们输出的单倍型的开关和翻转错误率是有利的,当比较它们与最先进的统计相位器。
The human genome is diploid, which requires assigning heterozygous single nucleotide polymorphisms (SNPs) to the two copies of the genome. The resulting haplotypes, lists of SNPs belonging to each copy, are crucial for downstream analyses in population genetics. Currently, statistical approaches, which are oblivious to direct read information, constitute the state-of-the-art. Haplotype assembly, which addresses phasing directly from sequencing reads, suffers from the fact that sequencing reads of the current generation are too short to serve the purposes of genome-wide phasing. While future-technology sequencing reads will contain sufficient amounts of SNPs per read for phasing, they are also likely to suffer from higher sequencing error rates. Currently, no haplotype assembly approaches exist that allow for taking both increasing read length and sequencing error information into account. Here, we suggest WhatsHap, the first approach that yields provably optimal solutions to the weighted minimum error correction problem in runtime linear in the number of SNPs. WhatsHap is a fixed parameter tractable (FPT) approach with coverage as the parameter. We demonstrate that WhatsHap can handle datasets of coverage up to 20x, and that 15x are generally enough for reliably phasing long reads, even at significantly elevated sequencing error rates. We also find that the switch and flip error rates of the haplotypes we output are favorable when comparing them with state-of-the-art statistical phasers.