WITCH: Improved Multiple Sequence Alignment Through Weighted Consensus Hidden Markov Model Alignment

WITCH: Improved Multiple Sequence Alignment Through Weighted Consensus Hidden Markov Model Alignment
复制标题

DOI:
10.1089/cmb.2021.0585
复制
发表时间:
2022-05-17
影响因子:
1.7
通讯作者:
Warnow, Tandy
Warnow, Tandy
中科院分区:
生物学4区
文献类型:
--
作者:
Shen, Chengze;Park, Minhyuk;Warnow, Tandy

文献摘要

被引文献

相似文献

准确的多序列比对在许多数据集上是具有挑战性的,包括那些大的、在高进化速率下进化的、或具有序列长度异质性的数据集。虽然在过去十年中在解决前两个挑战方面取得了实质性进展,但对于许多数据集来说,序列长度异质性仍然是一个重要问题。序列长度的异质性是由生物和技术原因引起的,包括在与序列相关的进化史中发生的大量插入或缺失(INDel),或包含未完全组装的序列。使用系统发育感知轮廓(UPP)的超大比对(Nguyen等人2015)是用于对表现出序列长度异质性的数据集进行比对的最准确的方法之一:它在其认为是“全长”的序列子集上构建比对,使用隐马尔可夫模型(HMM)的集合来表示该“主干比对”,然后基于从该集合中为该序列选择的HMM将每个剩余序列添加到主干比对中。我们的新方法加权一致HMM对齐(WITCH)在三个重要方面对UPP进行了改进:第一,它使用统计原理的技术来对HMM进行加权和排序;第二,它使用集成中的k&>1个HMM而不是单个HMM;第三,它使用考虑权重的一致算法来组合每个所选HMM的对齐。我们表明,与UPP和其他领先的比对方法相比,该方法提供了更高的比对精度,并且基于这些比对的最大似然树的精度也得到了提高。
Accurate multiple sequence alignment is challenging on many data sets, including those that are large, evolve under high rates of evolution, or have sequence length heterogeneity. While substantial progress has been made over the last decade in addressing the first two challenges, sequence length heterogeneity remains a significant issue for many data sets. Sequence length heterogeneity occurs for biological and technological reasons, including large insertions or deletions (indels) that occurred in the evolutionary history relating the sequences, or the inclusion of sequences that are not fully assembled. Ultra-large alignments using Phylogeny-Aware Profiles (UPP) (Nguyen et al. 2015) is one of the most accurate approaches for aligning data sets that exhibit sequence length heterogeneity: it constructs an alignment on the subset of sequences it considers "full-length, " represents this "backbone alignment " using an ensemble of hidden Markov models (HMMs), and then adds each remaining sequence into the backbone alignment based on an HMM selected for that sequence from the ensemble. Our new method, WeIghTed Consensus Hmm alignment (WITCH), improves on UPP in three important ways: first, it uses a statistically principled technique to weight and rank the HMMs; second, it uses k > 1 HMMs from the ensemble rather than a single HMM; and third, it combines the alignments for each of the selected HMMs using a consensus algorithm that takes the weights into account. We show that this approach provides improved alignment accuracy compared with UPP and other leading alignment methods, as well as improved accuracy for maximum likelihood trees based on these alignments.