Development of an unbiased statistical method for the analysis of unigenic evolution

Development of an unbiased statistical method for the analysis of unigenic evolution
复制标题

DOI:
10.1186/1471-2105-7-150
复制
发表时间:
2006-03-17
期刊:
影响因子:
3
通讯作者:
Wahl, LM
Wahl, LM
中科院分区:
生物学4区
文献类型:
--
作者:
Behrsin, CD;Brandl, CJ;Wahl, LM

文献摘要

被引文献

相似文献

背景:单基因进化是一种强大的遗传策略,涉及单个基因产物的随机突变,以描绘蛋白质的功能重要结构域。该方法包括选择保留功能的蛋白质变体,然后进行统计分析,比较每个残基的预期和观察到的突变频率。所得的每个残基的易变性指数在指定的密码子窗口上平均,以确定蛋白质的低易变性区域。正如最初描述的那样,对平均窗口长度变化的影响并没有完全阐明。此外,尚不清楚何时检查了足够的功能变体,以得出所有变体中保守的残基具有重要功能作用的结论。结果:我们证明了平均窗口的长度显著影响单个次变区域的识别和区域边界的划定。因此,我们设计了一个区域无关的卡方分析,消除了在窗口平均期间产生的信息损失,并消除了窗口长度的任意分配。我们还提出了一种方法来估计保守残基不被偶然突变的概率。此外,我们描述了期望突变频率的改进估计。结论:总的来说,这些方法大大扩展了现有方法对单基因进化数据的分析,从而能够全面、公正地鉴定对蛋白质功能至关重要的结构域,甚至可能是单个残基。
Background: Unigenic evolution is a powerful genetic strategy involving random mutagenesis of a single gene product to delineate functionally important domains of a protein. This method involves selection of variants of the protein which retain function, followed by statistical analysis comparing expected and observed mutation frequencies of each residue. Resultant mutability indices for each residue are averaged across a specified window of codons to identify hypomutable regions of the protein. As originally described, the effect of changes to the length of this averaging window was not fully eludicated. In addition, it was unclear when sufficient functional variants had been examined to conclude that residues conserved in all variants have important functional roles.Results: We demonstrate that the length of averaging window dramatically affects identification of individual hypomutable regions and delineation of region boundaries. Accordingly, we devised a region-independent chi-square analysis that eliminates loss of information incurred during window averaging and removes the arbitrary assignment of window length. We also present a method to estimate the probability that conserved residues have not been mutated simply by chance. In addition, we describe an improved estimation of the expected mutation frequency.Conclusion: Overall, these methods significantly extend the analysis of unigenic evolution data over existing methods to allow comprehensive, unbiased identification of domains and possibly even individual residues that are essential for protein function.