Composition-based statistics and translated nucleotide searches: improving the TBLASTN module of BLAST.

Composition-based statistics and translated nucleotide searches: improving the TBLASTN module of BLAST.
复制标题

DOI:
10.1186/1741-7007-4-41
复制
发表时间:
2006-12-07
期刊:
影响因子:
5.4
通讯作者:
Altschul SF
Altschul SF
中科院分区:
生物学2区
文献类型:
--
作者:
Gertz EM;Yu YK;Agarwala R;Schäffer AA;Altschul SF

文献摘要

参考文献

被引文献

相似文献

TBLASTN 是 BLAST 的一种操作模式,它将蛋白质序列与所有六个框架中翻译的核苷酸数据库进行比对。我们首次描述了 TBLASTN 的现代实现,重点关注用于实现翻译核苷酸搜索的基于组成的统计的新技术。基于组合的统计使用被比对的序列的组合来生成更准确的 E 值,从而可以更准确地区分真匹配和假匹配。直到最近,基于成分的统计数据仅可用于蛋白质-蛋白质搜索。它们现在可作为 TBLASTN 最新版本的命令行选项以及 NCBI BLAST Web 服务器上的 TBLASTN 选项。我们评估了 TBLASTN 的基线版本和使用不同类型的基于成分的统计数据的两个变体报告的 E 值的统计和检索准确性。为了测试 TBLASTN 的统计准确性,我们使用小鼠基因组中的乱序蛋白质和人类染色体数据库进行了 1000 次搜索。为了测试检索准确性,我们对以前用于评估蛋白质-蛋白质搜索的检索准确性的测试集进行了现代化改造并使其适应翻译搜索。我们表明,基于组合的统计极大地提高了 TBLASTN 的统计准确性,而检索准确性的代价很小。 TBLASTN 被广泛使用,因为人们通常希望将蛋白质与染色体或 mRNA 文库进行比较。基于成分的统计提高了 TBLASTN 结果的统计准确性,从而提高了可靠性。 TBLASTN 使用的算法并不广为人知,这里报告了一些最重要的算法。用于测试 TBLASTN 的数据可供下载,并且可能对翻译搜索算法的其他研究有用。
TBLASTN is a mode of operation for BLAST that aligns protein sequences to a nucleotide database translated in all six frames. We present the first description of the modern implementation of TBLASTN, focusing on new techniques that were used to implement composition-based statistics for translated nucleotide searches. Composition-based statistics use the composition of the sequences being aligned to generate more accurate E-values, which allows for a more accurate distinction between true and false matches. Until recently, composition-based statistics were available only for protein-protein searches. They are now available as a command line option for recent versions of TBLASTN and as an option for TBLASTN on the NCBI BLAST web server. We evaluate the statistical and retrieval accuracy of the E-values reported by a baseline version of TBLASTN and by two variants that use different types of composition-based statistics. To test the statistical accuracy of TBLASTN, we ran 1000 searches using scrambled proteins from the mouse genome and a database of human chromosomes. To test retrieval accuracy, we modernize and adapt to translated searches a test set previously used to evaluate the retrieval accuracy of protein-protein searches. We show that composition-based statistics greatly improve the statistical accuracy of TBLASTN, at a small cost to the retrieval accuracy. TBLASTN is widely used, as it is common to wish to compare proteins to chromosomes or to libraries of mRNAs. Composition-based statistics improve the statistical accuracy, and therefore the reliability, of TBLASTN results. The algorithms used by TBLASTN are not widely known, and some of the most important are reported here. The data used to test TBLASTN are available for download and may be useful in other studies of translated search algorithms.
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1016/s0097-8485(96)80004-0
发表时间: 1996-03-01
期刊: COMPUTERS & CHEMISTRY
影响因子: --
作者:
Gribskov, M;Robinson, NL
通讯作者: Robinson, NL
DOI: 10.1093/bioinformatics/16.3.190
发表时间: 2000-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Gotoh, O
通讯作者: Gotoh, O
DOI: 10.1006/jtbi.1994.1062
发表时间: 1994-03-21
影响因子: 2
作者:
HEIN, J
通讯作者: HEIN, J
DOI: 10.1038/ng0393-266
发表时间: 1993-03-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
GISH, W;STATES, DJ
通讯作者: STATES, DJ