Phylogenetic estimation of context-dependent substitution rates by maximum likelihood

Phylogenetic estimation of context-dependent substitution rates by maximum likelihood
复制标题

DOI:
10.1093/molbev/msh039
复制
发表时间:
2004-03-01
影响因子:
10.7
通讯作者:
Haussler, D
Haussler, D
中科院分区:
生物学1区
文献类型:
--
作者:
Siepel, A;Haussler, D

文献摘要

被引文献

相似文献

在编码区和非编码区的核苷酸取代是上下文依赖性的,在这个意义上,取代率取决于相邻碱基的身份。上下文相关的替代已经在两个序列和无根系统发育树的情况下建模,但它只以有限的方式与更一般的同源性相适应。在这篇文章中,扩展标准的系统发育模型,允许更好地处理上下文相关的替代,但仍然允许在合理的计算成本精确的推断。新模型大大提高了编码和非编码数据的拟合优度。考虑上下文的依赖性导致更大的改进比使用更丰富的替代模型或允许跨站点的速率变化,在站点独立性的假设下。所观察到的改进似乎来自三个独立的属性的模型:其明确的上下文相关的N-元组内的相邻站点,其能力,以适应重叠的N-元组的取代,其丰富的参数化的替代过程的表征。参数估计是使用期望最大化算法,与拟牛顿算法的最大化步骤,这种方法被证明是最好的普通牛顿方法参数丰富的模型。重叠元组是有效地处理假设马尔可夫依赖的观察基地在每个网站上的N - 1前的网站,和所需的。条件概率是用Felsenstein算法的扩展来计算的。基于哺乳动物基因组中约160,000个非编码位点的数据集估计的取代率表明明显的CpG效应,但它们也表明了复杂的上下文依赖性取代的总体模式,包括各种微妙的效应。基于编码区约300万个位点的估计表明,氨基酸取代率可以在核苷酸水平上学习,并表明密码子边界的背景效应是显着的。
Nucleotide substitution in both coding and noncoding regions is context-dependent, in the sense that substitution rates depend on the identity of neighboring bases. Context-dependent substitution has been modeled in the case of two sequences and an unrooted phylogenetic tree, but it has only been accommodated in limited ways with more general phylogenies. In this article, extensions are presented to standard phylogenetic models that allow for better handling of context-dependent substitution, yet still permit exact inference at reasonable computational cost. The new models improve goodness of fit substantially for both coding and noncoding data. Considering context dependence leads to much larger improvements than does using a richer substitution model or allowing for rate variation across sites, under the assumption of site independence. The observed improvements appear to derive from three separate properties of the models: their explicit characterization of context-dependent substitution within N-tuples of adjacent sites, their ability to accommodate overlapping N-tuples, and their rich parameterization of the substitution process. Parameter estimation is accomplished using an expectation maximization algorithm, with a quasi-Newton algorithm for the maximization step; this approach is shown to be preferable to ordinary Newton methods for parameter-rich models. Overlapping tuples are efficiently handled by assuming Markov dependence of the observed bases at each site on those at the N - 1 preceding sites, and the required. conditional probabilities are computed with an extension of Felsenstein's algorithm. Estimated substitution rates based on a data set of about 160,000 noncoding sites in mammalian genomes indicate a pronounced CpG effect, but they also suggest a complex overall pattern of context-dependent substitution, comprising a variety of subtle effects. Estimates based on about 3 million sites in coding regions demonstrate that amino acid substitution rates can be learned at the nucleotide level, and suggest that context effects across codon boundaries are significant.