A Dirichlet process model for detecting positive selection in protein-coding DNA sequences

A Dirichlet process model for detecting positive selection in protein-coding DNA sequences
复制标题

DOI:
10.1073/pnas.0508279103
复制
发表时间:
2006-04-18
影响因子:
11.1
通讯作者:
Pond, SLK
Pond, SLK
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Huelsenbeck, JP;Jain, S;Pond, SLK

文献摘要

被引文献

相似文献

大多数在分子水平上检测达尔文自然选择的方法依赖于估计蛋白质编码DNA序列比对中非同义和同义变化的速率或数量。在这些方法中的一些中,允许非同义取代率在整个序列中变化,从而允许鉴定正自然选择下的单个氨基酸位置。然而,目前还不清楚应该使用哪种概率分布来描述非同义取代率在整个序列中的变化。一种广泛使用的解决方案是将序列中非同义速率的变化建模为几个离散或连续概率分布的混合。不幸的是,很少有群体遗传学理论告诉我们适当的概率分布的非同义取代率的位点变异。在这里,我们描述了一种方法,通过使用狄利克雷过程混合模型的非同义替代率的变化建模。狄利克雷过程允许存在可计数的无限数量的非同义速率类,并且在适应非同义替代率的不同潜在分布方面非常灵活。我们采用完全贝叶斯方法实现了该模型,模型的所有参数都被视为随机变量。
Most methods for detecting Darwinian natural selection at the molecular level rely on estimating the rates or numbers of non-synonymous and synonymous changes in an alignment of protein-coding DNA sequences. In some of these methods, the nonsynonymous rate of substitution is allowed to vary across the sequence, permitting the identification of single amino acid positions that are under positive natural selection. However, it is unclear which probability distribution should be used to describe how the nonsynonymous rate of substitution varies across the sequence. One widely used solution is to model variation in the nonsynonymous rate across the sequence as a mixture of several discrete or continuous probability distributions. Unfortunately, there is little population genetics theory to inform us of the appropriate probability distribution for among-site variation in the nonsynonymous rate of substitution. Here, we describe an approach to modeling variation in the nonsynonymous rate of substitution by using a Dirichlet process mixture model. The Dirichlet process allows there to be a countably infinite number of nonsynonymous rate classes and is very flexible in accommodating different potential distributions for the nonsynonymous rate of substitution. We implemented the model in a fully Bayesian approach, with all parameters of the model considered as random variables.