Measuring global credibility with application to local sequence alignment.

Measuring global credibility with application to local sequence alignment.
复制标题

DOI:
10.1371/journal.pcbi.1000077
复制
发表时间:
2008-05-16
影响因子:
4.3
通讯作者:
Lawrence CE
Lawrence CE
中科院分区:
生物学2区
文献类型:
--
作者:
Webb-Robertson BJ;McCue LA;Lawrence CE

文献摘要

参考文献

被引文献

相似文献

计算生物学充满了高维(high-D)离散预测和推理问题,包括序列比对、RNA结构预测、系统发育推断、基序发现、路径预测和统计遗传学中的模型选择问题。尽管在这些情况下的预测和推断是不确定的,但很少有人把注意力集中在制定全球不确定性措施上。不管使用什么过程来产生预测,当过程提供单个答案时,该答案是从解决方案集合(所有可能解决方案的集合)中选择的一个点估计。对于高d离散空间,这些集合是巨大的,因此有相当大的不确定性。我们建议使用贝叶斯可信度限来描述这种不确定性,其中(1−α)%, 0≤α≤1,可信度限是包含(1−α)%后验分布的超球的最小汉明距离半径。因为序列比对可以说是计算生物学中最广泛使用的程序,我们在这里使用它来使这些一般概念更加具体。最大相似性估计器(即,最大化似然的对齐)和质心估计器(即,最小化从后验加权对齐集合的平均汉明距离的对齐)用于演示贝叶斯可信度限制对对齐估计器的应用。对来自6种希瓦氏菌的20对人/啮齿动物同源序列和125对同源序列比对的贝叶斯信度限应用表明,这些物种的启动子序列比对的信度限差异很大,质心比对比传统的最大相似性比对具有更严格的信度限。序列比对是许多计算生物学应用所使用的基础能力,如系统发育重建和共同调节机制的识别。序列比对方法通常寻求一对序列之间的高分比对,并为这一单一比对分配统计显著性。然而,由于两个(或更多)序列的单一比对是一个点估计,它可能不能代表这些序列的可能比对的整个集合(集合);因此,在巨大的可能性集合中,任何一种排列都可能存在相当大的不确定性。为了解决提议的对齐的不确定性,我们使用贝叶斯概率方法来评估在可能对齐的整个集合的上下文中对齐的可靠性。我们的方法对集合的成员偏离选定的组合的程度进行全局评估,从而确定可信度限制。在对流行的最大相似度对齐和质心对齐(即对齐后验分布中心的对齐)的评估中,我们发现质心比最大相似度对齐产生更严格的可信度限制(平均而言)。除了通常对点估计的误差限制感兴趣之外,我们对校准可信度限制的实质性变化的发现表明,更广泛地采用这些限制,因此在后续使用校准之前描述了误差程度。
Computational biology is replete with high-dimensional (high-D) discrete prediction and inference problems, including sequence alignment, RNA structure prediction, phylogenetic inference, motif finding, prediction of pathways, and model selection problems in statistical genetics. Even though prediction and inference in these settings are uncertain, little attention has been focused on the development of global measures of uncertainty. Regardless of the procedure employed to produce a prediction, when a procedure delivers a single answer, that answer is a point estimate selected from the solution ensemble, the set of all possible solutions. For high-D discrete space, these ensembles are immense, and thus there is considerable uncertainty. We recommend the use of Bayesian credibility limits to describe this uncertainty, where a (1−α)%, 0≤α≤1, credibility limit is the minimum Hamming distance radius of a hyper-sphere containing (1−α)% of the posterior distribution. Because sequence alignment is arguably the most extensively used procedure in computational biology, we employ it here to make these general concepts more concrete. The maximum similarity estimator (i.e., the alignment that maximizes the likelihood) and the centroid estimator (i.e., the alignment that minimizes the mean Hamming distance from the posterior weighted ensemble of alignments) are used to demonstrate the application of Bayesian credibility limits to alignment estimators. Application of Bayesian credibility limits to the alignment of 20 human/rodent orthologous sequence pairs and 125 orthologous sequence pairs from six Shewanella species shows that credibility limits of the alignments of promoter sequences of these species vary widely, and that centroid alignments dependably have tighter credibility limits than traditional maximum similarity alignments. Sequence alignment is the cornerstone capability used by a multitude of computational biology applications, such as phylogeny reconstruction and identification of common regulatory mechanisms. Sequence alignment methods typically seek a high-scoring alignment between a pair of sequences, and assign a statistical significance to this single alignment. However, because a single alignment of two (or more) sequences is a point estimate, it may not be representative of the entire set (ensemble) of possible alignments of those sequences; thus, there may be considerable uncertainty associated with any one alignment among an immense ensemble of possibilities. To address the uncertainty of a proposed alignment, we used a Bayesian probabilistic approach to assess an alignment's reliability in the context of the entire ensemble of possible alignments. Our approach performs a global assessment of the degree to which the members of the ensemble depart from a selected alignment, thereby determining a credibility limit. In an evaluation of the popular maximum similarity alignment and the centroid alignment (i.e., the alignment that is in the center of the posterior distribution of alignments), we find that the centroid yields tighter credibility limits (on average) than the maximum similarity alignment. Beyond the usual interest in putting error limits on point estimates, our findings of substantial variability in credibility limits of alignments argue for wider adoption of these limits, so the degree of error is delineated prior to the subsequent use of the alignments.
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1089/cmb.1998.5.493
发表时间: 1998-09-01
影响因子: 1.7
作者:
Holmes, I;Durbin, R
通讯作者: Durbin, R
DOI: 10.1093/protein/8.10.999
发表时间: 1995-10-01
期刊: PROTEIN ENGINEERING
影响因子: --
作者:
Miyazawa, S
通讯作者: Miyazawa, S
DOI: 10.1002/pro.5560040613
发表时间: 1995-06-01
期刊: PROTEIN SCIENCE
影响因子: 8
作者:
PEARSON, WR
通讯作者: PEARSON, WR
DOI: 10.1073/pnas.85.8.2444
发表时间: 1988-04-01
影响因子: 11.1
作者:
PEARSON, WR;LIPMAN, DJ
通讯作者: LIPMAN, DJ