Learning to Count: Robust Estimates for Labeled Distances between Molecular Sequences

Learning to Count: Robust Estimates for Labeled Distances between Molecular Sequences
复制标题

DOI:
10.1093/molbev/msp003
复制
发表时间:
2009-04-01
影响因子:
10.7
通讯作者:
Suchard, Marc A.
Suchard, Marc A.
中科院分区:
生物学1区
文献类型:
--
作者:
O'Brien, John D.;Minin, Vladimir N.;Suchard, Marc A.

文献摘要

被引文献

相似文献

研究人员通常使用连续时间马尔可夫链模型来估计分子序列之间的距离。我们提出了一种新的方法,稳健计数,以防止可能严重的偏差产生的模型规格错误。我们通过推广传统的距离估计来实现这种鲁棒性,以结合观察到的成对序列比对中发现的站点模式的经验分布。我们灵活的框架允许仅基于可能替换的子集来计算距离。由此,我们展示了如何估计标记密码子距离,如同义或非同义替换的预期数量。我们提出了两个模拟研究。首先比较了传统和稳健标记核苷酸估计器的相对偏差和方差。在第二个模拟中,我们证明了鲁棒计数仅基于易于拟合的核苷酸取代模型提供准确的同义和非同义距离估计,而不需要计算昂贵的密码子模型。我们用三个实证例子来总结。在前两个例子中,我们利用标记密码子距离研究了甲型流感血凝素基因的进化动力学。在最后的例子中,我们展示了使用鲁棒同义距离来减轻收敛进化对HIV传播网络系统发育分析的影响的优势。
Researchers routinely estimate distances between molecular sequences using continuous-time Markov chain models. We present a new method, robust counting, that protects against the possibly severe bias arising from model misspecification. We achieve this robustness by generalizing the conventional distance estimation to incorporate the empirical distribution of site patterns found in the observed pairwise sequence alignment. Our flexible framework allows for computing distances based only on a subset of possible substitutions. From this, we show how to estimate labeled codon distances, such as expected numbers of synonymous or nonsynonymous substitutions. We present two simulation studies. The first compares the relative bias and variance of conventional and robust labeled nucleotide estimators. In the second simulation, we demonstrate that robust counting furnishes accurate synonymous and nonsynonymous distance estimates based only on easy-to-fit models of nucleotide substitution, bypassing the need for computationally expensive codon models. We conclude with three empirical examples. In the first two examples, we investigate the evolutionary dynamics of the influenza A hemagglutinin gene using labeled codon distances. In the final example, we demonstrate the advantages of using robust synonymous distances to alleviate the effect of convergent evolution on phylogenetic analysis of an HIV transmission network.