MAGUS+eHMMs: improved multiple sequence alignment accuracy for fragmentary sequences.

MAGUS+eHMMs: improved multiple sequence alignment accuracy for fragmentary sequences.
复制标题

DOI:
10.1093/bioinformatics/btab788
复制
发表时间:
2022-01-27
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Warnow T
Warnow T
中科院分区:
其他
文献类型:
--
作者:
Shen C;Zaharias P;Warnow T

文献摘要

参考文献

被引文献

相似文献

多序列比对是许多生物信息学流程中的第一步,包括系统发育估计、蛋白质结构预测和扩增子或宏基因组数据集中产生的读数的分类学识别等。然而,对于表现出显着序列长度异质性的数据集,比对估计具有挑战性,特别是当数据集由于包含由下一代测序技术生成的读数或重叠群而具有片段序列时。在这里,我们研究了当数据集包含大量片段序列时改进比对估计的技术。我们发现 MAGUS(一种最近开发的 MSA 方法)在许多条件下对片段序列相当稳健,并且使用两阶段方法,其中 MAGUS 用于比对选定的“主干序列”,并使用隐马尔可夫模型集合将剩余序列添加到比对中,进一步提高了比对准确性。 MAGUS 与 eHMM 集成(即 MAGUS+eHMM)的结合明显改进了 UPP,UPP 是之前用于对齐具有高水平碎片的数据集的领先方法。 UPP 可在 https://github.com/smirarab/sepp 上获取,MAGUS 可在 https://github.com/vlasmirnov/MAGUS 上获取。 MAGUS+eHMM 可以通过运行 MAGUS 来获得骨干对齐,然后使用骨干对齐作为 UPP 的输入来执行。 补充数据可在生物信息学在线获取。
Multiple sequence alignment is an initial step in many bioinformatics pipelines, including phylogeny estimation, protein structure prediction and taxonomic identification of reads produced in amplicon or metagenomic datasets, etc. Yet, alignment estimation is challenging on datasets that exhibit substantial sequence length heterogeneity, and especially when the datasets have fragmentary sequences as a result of including reads or contigs generated by next-generation sequencing technologies. Here, we examine techniques that have been developed to improve alignment estimation when datasets contain substantial numbers of fragmentary sequences. We find that MAGUS, a recently developed MSA method, is fairly robust to fragmentary sequences under many conditions, and that using a two-stage approach where MAGUS is used to align selected ‘backbone sequences’ and the remaining sequences are added into the alignment using ensembles of Hidden Markov Models further improves alignment accuracy. The combination of MAGUS with the ensemble of eHMMs (i.e. MAGUS+eHMMs) clearly improves on UPP, the previous leading method for aligning datasets with high levels of fragmentation. UPP is available on https://github.com/smirarab/sepp, and MAGUS is available on https://github.com/vlasmirnov/MAGUS. MAGUS+eHMMs can be performed by running MAGUS to obtain the backbone alignment, and then using the backbone alignment as an input to UPP. Supplementary data are available at Bioinformatics online.
DOI: 10.1093/bioinformatics/btab023
发表时间: 2021-07-27
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Shah N;Molloy EK;Pop M;Warnow T
通讯作者: Warnow T
DOI: 10.1093/bioinformatics/btr553
发表时间: 2011-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Mirarab, Siavash;Warnow, Tandy
通讯作者: Warnow, Tandy
DOI: 10.1093/bioinformatics/14.2.157
发表时间: 1998-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Stoye, J;Evers, D;Meyer, F
通讯作者: Meyer, F
DOI: 10.1093/sysbio/syaa058
发表时间: 2021-02-10
期刊: Systematic biology
影响因子: 6.5
作者:
Smirnov V;Warnow T
通讯作者: Warnow T
DOI: 10.1186/1471-2105-7-471
发表时间: 2006-10-24
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Nuin, Paulo A S;Wang, Zhouzhi;Tillier, Elisabeth R M
通讯作者: Tillier, Elisabeth R M