Sensitive and error-tolerant annotation of protein-coding DNA with BATH.

Sensitive and error-tolerant annotation of protein-coding DNA with BATH.
复制标题

使用 BATH 对蛋白质编码 DNA 进行灵敏且容错的注释。

DOI:
10.1101/2023.12.31.573773
复制
发表时间:
2024
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Wheeler,TravisJ
Wheeler,TravisJ
中科院分区:
--
文献类型:
--
作者:
Krause,GenevieveR;Shands,Walt;Wheeler,TravisJ

文献摘要

相似文献

摘要我们介绍了Bath,这是一个用于高度敏感地注释蛋白质编码DNA的工具,该工具基于DNA与蛋白质序列数据库或轮廓隐藏马尔可夫模型(PHMM)的直接比对。Bath构建在HMMER3代码库之上,通过提供直观的输入接口和易于解释的输出,简化了基于PHMM的翻译序列批注的批注工作流程。巴斯还引入了新的移码感知算法来检测导致移码的核苷酸插入和缺失(INDELs)。Bath在注释不含错误的序列方面与HMMER3的准确性相匹配,并且比所有用于注释包含核苷酸INDELs的序列的测试工具产生更高的准确性。这些结果表明,当需要很高的注释敏感性时,尤其是当移码错误预计会中断蛋白质编码区时,应该使用Bath,就像长读测序数据和假基因的情况一样。可用性和实施该软件可在https://github.com/TravisWheelerLab/BATH.获得
SummaryWe present BATH, a tool for highly sensitive annotation of protein-coding DNA based on direct alignment of that DNA to a database of protein sequences or profile hidden Markov models (pHMMs). BATH is built on top of the HMMER3 code base, and simplifies the annotation workflow for pHMM-based translated sequence annotation by providing a straightforward input interface and easy-to-interpret output. BATH also introduces novel frameshift-aware algorithms to detect frameshift-inducing nucleotide insertions and deletions (indels). BATH matches the accuracy of HMMER3 for annotation of sequences containing no errors, and produces superior accuracy to all tested tools for annotation of sequences containing nucleotide indels. These results suggest that BATH should be used when high annotation sensitivity is required, particularly when frameshift errors are expected to interrupt protein-coding regions, as is true with long-read sequencing data and in the context of pseudogenes.Availability and implementationThe software is available at https://github.com/TravisWheelerLab/BATH.