IDENTIFICATION OF PROTEIN CODING REGIONS BY DATABASE SIMILARITY SEARCH

IDENTIFICATION OF PROTEIN CODING REGIONS BY DATABASE SIMILARITY SEARCH
复制标题

DOI:
10.1038/ng0393-266
复制
发表时间:
1993-03-01
期刊:
影响因子:
30.8
通讯作者:
STATES, DJ
STATES, DJ
中科院分区:
生物学1区
文献类型:
--
作者:
GISH, W;STATES, DJ

文献摘要

被引文献

相似文献

翻译的核苷酸序列和已知的生物蛋白质之间的序列相似性可以为同源编码区的存在提供强有力的证据,即使在远亲基因之间也是如此。计算机程序BLASTX在一个程序步骤中进行核苷酸查询序列的概念翻译,然后进行蛋白质数据库搜索。我们表征了BLASTX识别对查询序列中存在的取代、插入和缺失错误以及对序列趋异的敏感性。阅读帧在存在1%查询错误的情况下被可靠地鉴定,这是一级序列数据的典型比率。BLASTX适用于在数据最容易包含错误的情况下尽早用于中等规模和大规模测序项目。
Sequence similarity between a translated nucleotide sequence and a known biological protein can provide strong evidence for the presence of a homologous coding region, even between distantly related genes. The computer program BLASTX performed conceptual translation of a nucleotide query sequence followed by a protein database search in one programmatic step. We characterized the sensitivity of BLASTX recognition to the presence of substitution, insertion and deletion errors in the query sequence and to sequence divergence. Reading frames were reliably identified in the presence of 1 % query errors, a rate that is typical for primary sequence data. BLASTX is appropriate for use in moderate and large scale sequencing projects at the earliest opportunity, when the data are most prone to containing errors.