Haplotype-aware diplotyping from noisy long reads

Haplotype-aware diplotyping from noisy long reads
复制标题

DOI:
10.1186/s13059-019-1709-0
复制
发表时间:
2019-06-03
期刊:
影响因子:
12.3
通讯作者:
Paten, Benedict
Paten, Benedict
中科院分区:
生物学1区
文献类型:
--
作者:
Ebler, Jana;Haukness, Marina;Paten, Benedict

文献摘要

被引文献

相似文献

目前单核苷酸变异的基因分型方法依赖于第二代测序设备的短而准确的读数。目前,第三代测序平台正在迅速普及,但缺乏利用其长但容易出错的读数进行基因分型的方法。在这里,我们引入了一种新颖的统计框架,用于从嘈杂的长读取中联合推断单倍型和基因型,我们称之为二倍型。我们的技术充分利用长读提供的链接信息。我们验证了数十万个尚未包含在瓶中基因组工作的高可信度参考集中的候选变体。
Current genotyping approaches for single-nucleotide variations rely on short, accurate reads from second-generation sequencing devices. Presently, third-generation sequencing platforms are rapidly becoming more widespread, yet approaches for leveraging their long but error-prone reads for genotyping are lacking. Here, we introduce a novel statistical framework for the joint inference of haplotypes and genotypes from noisy long reads, which we term diplotyping. Our technique takes full advantage of linkage information provided by long reads. We validate hundreds of thousands of candidate variants that have not yet been included in the high-confidence reference set of the Genome-in-a-Bottle effort.