P-smoother: efficient PBWT smoothing of large haplotype panels.

P-smoother: efficient PBWT smoothing of large haplotype panels.
复制标题

DOI:
10.1093/bioadv/vbac045
复制
发表时间:
2022
期刊:
Bioinformatics advances
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

随着大的单倍型面板变得越来越可用,有效的字符串匹配算法,如位置Burrows-Wheeler变换(PBWT)是有前途的识别共享单倍型。然而,最近的突变和基因分型错误造成偶尔的错配,提出了精确的单倍型匹配的挑战。以前的解决方案是基于概率模型或种子和扩展算法,被动地容忍不匹配。在这里,我们提出了一个基于PBWT的平滑算法,P平滑,积极“纠正”这些不匹配,从而“平滑”面板。P-smoother运行基于PBWT的双向面板扫描,基于整体单倍型匹配背景翻转不匹配的等位基因,我们称之为IBD(血统相同)先验。在一个有4000个单倍型和0.2%错误率的模拟面板中,我们表明它可以可靠地纠正85%的错误。因此,PBWT算法在平滑面板上运行可以识别比在未平滑面板上更多的成对IBD段。最引人注目的是,在平滑面板上运行的PBWT聚类算法,我们称之为PS聚类,实现了识别多路IBD段的最先进性能,这是计算界多年来的一个挑战性问题。我们还表明,PS集群是足够有效的英国生物银行数据。因此,P-smoother为生物库规模的单倍型面板的有效容错算法开辟了新的可能性。源代码可在github.com/ZhiGroup/P-smoother上获得。
As large haplotype panels become increasingly available, efficient string matching algorithms such as positional Burrows-Wheeler transformation (PBWT) are promising for identifying shared haplotypes. However, recent mutations and genotyping errors create occasional mismatches, presenting challenges for exact haplotype matching. Previous solutions are based on probabilistic models or seed-and-extension algorithms that passively tolerate mismatches. Here, we propose a PBWT-based smoothing algorithm, P-smoother, to actively ‘correct’ these mismatches and thus ‘smooth’ the panel. P-smoother runs a bidirectional PBWT-based panel scanning that flips mismatching alleles based on the overall haplotype matching context, which we call the IBD (identical-by-descent) prior. In a simulated panel with 4000 haplotypes and a 0.2% error rate, we show it can reliably correct 85% of errors. As a result, PBWT algorithms running over the smoothed panel can identify more pairwise IBD segments than that over the unsmoothed panel. Most strikingly, a PBWT-cluster algorithm running over the smoothed panel, which we call PS-cluster, achieves state-of-the-art performance for identifying multiway IBD segments, a challenging problem in the computational community for years. We also showed that PS-cluster is adequately efficient for UK Biobank data. Therefore, P-smoother opens up new possibilities for efficient error-tolerating algorithms for biobank-scale haplotype panels. Source code is available at github.com/ZhiGroup/P-smoother.
DOI: 10.1371/journal.pcbi.1004842
发表时间: 2016-05
影响因子: 4.3
作者:
Kelleher J;Etheridge AM;McVean G
通讯作者: McVean G
DOI: 10.1016/j.ajhg.2011.04.023
发表时间: 2011-06-10
影响因子: 9.8
作者:
Gusev, Alexander;Kenny, Eimear E.;Pe'er, Itsik
通讯作者: Pe'er, Itsik
DOI: 10.1093/molbev/msaa328
发表时间: 2021-05-04
影响因子: 10.7
作者:
Freyman WA;McManus KF;Shringarpure SS;Jewett EM;Bryc K;23 and Me Research Team;Auton A
通讯作者: Auton A
DOI: 10.1038/ng.3679
发表时间: 2016-11
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Loh, Po-Ru;Danecek, Petr;Palamara, Pier Francesco;Fuchsberger, Christian;Reshef, Yakir A.;Finucane, Hilary K.;Schoenherr, Sebastian;Forer, Lukas;McCarthy, Shane;Abecasis, Goncalo R.;Durbin, Richard;Price, Alkes L.
通讯作者: Price, Alkes L.
DOI: 10.1101/gr.115360.110
发表时间: 2011-07-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Moltke, Ida;Albrechtsen, Anders;Nielsen, Rasmus
通讯作者: Nielsen, Rasmus