Ambiguity Coding Allows Accurate Inference of Evolutionary Parameters from Alignments in an Aggregated State-Space.

Ambiguity Coding Allows Accurate Inference of Evolutionary Parameters from Alignments in an Aggregated State-Space.
复制标题

模棱两可的编码允许从汇总状态空间中的比对准确推断进化参数。

DOI:
10.1093/sysbio/syaa036
复制
发表时间:
2021-01-01
期刊:
影响因子:
6.5
通讯作者:
Goldman N
Goldman N
中科院分区:
生物学1区
文献类型:
--
作者:
Weber CC;Perron U;Casey D;Yang Z;Goldman N

文献摘要

参考文献

被引文献

相似文献

我们如何才能最好地了解蛋白质进化的历史?理想情况下,一个序列进化的模型应该既能捕捉到产生遗传变异的过程,又能捕捉到决定哪些变化是固定的功能约束。然而,在实践中,最合适的方法可能只是一个结合了方便的输入数据的能力,返回有用的参数估计。例如,我们可能感兴趣的是选择强度的测量(通常使用密码子模型获得)或祖先结构(使用基于推断的氨基酸序列和侧链构型的结构建模获得)。但是,如果相关状态空间中的数据不容易获得呢?我们表明,它是可能的,以获得准确的估计,使用一个既定的方法处理缺失数据的输出。将对齐中观察到的字符编码为较大状态空间中字符的模糊表示,允许将具有所需特征的模型应用于缺乏通常所需分辨率的数据。这种策略是可行的,因为通过观察到的空间所采取的进化路径包含关于在“看不见的”状态空间中可能被访问的状态的信息。为了说明这一点,我们考虑两个例子与氨基酸序列作为输入。我们发现,一个参数描述的选择非同义和同义的变化的相对强度,可以估计在一个无偏的方式使用一个标准的61状态密码子模型的适应版本。使用模拟和经验数据,我们发现,祖先的氨基酸侧链构型可以推断通过应用55状态的经验模型,20状态的氨基酸数据。在可行的情况下,结合模糊编码和完全解析数据的输入可以提高准确性。与考虑完整旋转异构体状态信息的基准相比,将结构信息添加到氨基酸比对中少至12.5%的序列中导致显着的祖先重建性能。这些例子表明,我们的方法允许恢复的进化信息从序列中,它以前是不可访问的。[祖先重建;自然选择;蛋白质结构;状态空间;替代模型。]
How can we best learn the history of a protein’s evolution? Ideally, a model of sequence evolution should capture both the process that generates genetic variation and the functional constraints determining which changes are fixed. However, in practical terms the most suitable approach may simply be the one that combines the convenience of easily available input data with the ability to return useful parameter estimates. For example, we might be interested in a measure of the strength of selection (typically obtained using a codon model) or an ancestral structure (obtained using structural modeling based on inferred amino acid sequence and side chain configuration). But what if data in the relevant state-space are not readily available? We show that it is possible to obtain accurate estimates of the outputs of interest using an established method for handling missing data. Encoding observed characters in an alignment as ambiguous representations of characters in a larger state-space allows the application of models with the desired features to data that lack the resolution that is normally required. This strategy is viable because the evolutionary path taken through the observed space contains information about states that were likely visited in the “unseen” state-space. To illustrate this, we consider two examples with amino acid sequences as input. We show that , a parameter describing the relative strength of selection on nonsynonymous and synonymous changes, can be estimated in an unbiased manner using an adapted version of a standard 61-state codon model. Using simulated and empirical data, we find that ancestral amino acid side chain configuration can be inferred by applying a 55-state empirical model to 20-state amino acid data. Where feasible, combining inputs from both ambiguity-coded and fully resolved data improves accuracy. Adding structural information to as few as 12.5% of the sequences in an amino acid alignment results in remarkable ancestral reconstruction performance compared to a benchmark that considers the full rotamer state information. These examples show that our methods permit the recovery of evolutionary information from sequences where it has previously been inaccessible. [Ancestral reconstruction; natural selection; protein structure; state-spaces; substitution models.]
DOI: 10.1002/prot.22488
发表时间: 2009-12
影响因子: 2.9
作者:
Krivov, Georgii G.;Shapovalov, Maxim V.;Dunbrack, Roland L., Jr.
通讯作者: Dunbrack, Roland L., Jr.
DOI: 10.1093/nar/gky949
发表时间: 2019-01-08
影响因子: 14.9
作者:
wwPDB consortium
通讯作者: wwPDB consortium
DOI: 10.1007/bf02198858
发表时间: 1996-02-01
影响因子: 3.9
作者:
Koshi, JM;Goldstein, RA
通讯作者: Goldstein, RA
DOI: 10.1126/science.1138709
发表时间: 2007-04-13
期刊: SCIENCE
影响因子: 56.9
作者:
Schweitzer, Mary Higby;Suo, Zhiyong;Horner, John R.
通讯作者: Horner, John R.
DOI: 10.1093/molbev/msn067
发表时间: 2008-07-01
影响因子: 10.7
作者:
Le, Si Quang;Gascuel, Olivier
通讯作者: Gascuel, Olivier