The construction and use of log-odds substitution scores for multiple sequence alignment.

The construction and use of log-odds substitution scores for multiple sequence alignment.
复制标题

DOI:
10.1371/journal.pcbi.1000852
复制
发表时间:
2010-07-15
影响因子:
4.3
通讯作者:
Yu YK
Yu YK
中科院分区:
生物学2区
文献类型:
--
作者:
Altschul SF;Wootton JC;Zaslavsky E;Yu YK

文献摘要

参考文献

被引文献

相似文献

大多数成对和多序列比对程序寻求具有最佳得分的比对。定义这样的分数的核心是选择比对的氨基酸或核苷酸的一组取代分数。对于局部成对比对,替代评分隐含地具有对数优势形式。我们现在将对数几率形式主义扩展到多重比对,使用贝叶斯方法从描述相关字母列的先验分布构建“BILD”(“贝叶斯积分对数几率”)替代分数。这种方法以前仅用于定义将单个序列与序列谱进行比对的分数,但它具有更广泛的适用性。我们描述了如何有效地计算BILD分数,并说明了它们在Gibbs抽样优化程序,间隙对齐和隐马尔可夫模型配置文件的构建中的用途。BILD分数使得能够自动选择最佳基序和结构域模型宽度,并且可以告知是否在多重比对中包括序列的决定,以及插入和缺失位置的选择。其他应用包括将相关序列分类为亚家族,以及轮廓-轮廓比对分数的定义。尽管完全实现的多重比对程序必须依赖于多于取代评分,但是可以修改许多现有的多重比对程序以采用BILD评分。我们说明了如何简单的BILD评分为基础的策略可以提高识别的DNA结合域,包括在弓形虫和恶性疟原虫的Api-AP 2结构域。多序列比对是生物学研究的基本工具,广泛用于鉴定DNA或蛋白质分子的重要区域,推断其生物学功能,重建祖先,以及许多其他应用。序列比较程序的有效性和准确性关键取决于它们用于测量序列相似性的评分系统的质量。为了比较成对的DNA或蛋白质序列,构建相似性度量的最佳策略早已被理解,但关于如何测量多个(即两个以上)序列之间的相似性缺乏共识。在本文中,我们描述了一个自然的推广到多对齐的成对相似性的公认的措施。通过采用这种相似性度量,可以使用于比较和分析DNA或蛋白质分子或模拟蛋白质结构域家族的各种方法更加灵敏和精确。我们说明了我们的措施可以提高重要的DNA结合域的识别。
Most pairwise and multiple sequence alignment programs seek alignments with optimal scores. Central to defining such scores is selecting a set of substitution scores for aligned amino acids or nucleotides. For local pairwise alignment, substitution scores are implicitly of log-odds form. We now extend the log-odds formalism to multiple alignments, using Bayesian methods to construct “BILD” (“Bayesian Integral Log-odds”) substitution scores from prior distributions describing columns of related letters. This approach has been used previously only to define scores for aligning individual sequences to sequence profiles, but it has much broader applicability. We describe how to calculate BILD scores efficiently, and illustrate their uses in Gibbs sampling optimization procedures, gapped alignment, and the construction of hidden Markov model profiles. BILD scores enable automated selection of optimal motif and domain model widths, and can inform the decision of whether to include a sequence in a multiple alignment, and the selection of insertion and deletion locations. Other applications include the classification of related sequences into subfamilies, and the definition of profile-profile alignment scores. Although a fully realized multiple alignment program must rely upon more than substitution scores, many existing multiple alignment programs can be modified to employ BILD scores. We illustrate how simple BILD score based strategies can enhance the recognition of DNA binding domains, including the Api-AP2 domain in Toxoplasma gondii and Plasmodium falciparum. Multiple sequence alignment is a fundamental tool of biological research, widely used to identify important regions of DNA or protein molecules, to infer their biological functions, to reconstruct ancestries, and in numerous other applications. The effectiveness and accuracy of sequence comparison programs depends crucially upon the quality of the scoring systems they use to measure sequence similarity. To compare pairs of DNA or protein sequences, the best strategy for constructing similarity measures has long been understood, but there has been a lack of consensus about how to measure similarity among multiple (i.e. more than two) sequences. In this paper, we describe a natural generalization to multiple alignment of the accepted measure of pairwise similarity. A large variety of methods that are used to compare and analyze DNA or protein molecules, or to model protein domain families, could be rendered more sensitive and precise by adopting this similarity measure. We illustrate how our measure can enhance the recognition of important DNA binding domains.
DOI: 10.1016/s0022-5193(89)80196-1
发表时间: 1989-06-08
影响因子: 2
作者:
ALTSCHUL, SF
通讯作者: ALTSCHUL, SF
DOI: 10.1089/cmb.1995.2.9
发表时间: 1995-01-01
期刊: Journal of computational biology : a journal of computational molecular cell biology
影响因子: --
作者:
Eddy, S R;Mitchison, G;Durbin, R
通讯作者: Durbin, R
DOI: 10.1093/nar/gki709
发表时间: 2005
影响因子: 14.9
作者:
Balaji S;Babu MM;Iyer LM;Aravind L
通讯作者: Aravind L
DOI: 10.1016/0022-2836(86)90252-4
发表时间: 1986-09-20
影响因子: 5.6
作者:
BACON, DJ;ANDERSON, WF
通讯作者: ANDERSON, WF
DOI: 10.1093/bioinformatics/btg158
发表时间: 2003-07-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Edgar, RC;Sjölander, K
通讯作者: Sjölander, K