A novel approach to detecting and measuring recombination: New insights into evolution in viruses, bacteria, and mitochondria

A novel approach to detecting and measuring recombination: New insights into evolution in viruses, bacteria, and mitochondria
复制标题

DOI:
10.1093/oxfordjournals.molbev.a003928
复制
发表时间:
2001-08-01
影响因子:
10.7
通讯作者:
Worobey, M
Worobey, M
中科院分区:
生物学1区
文献类型:
--
作者:
Worobey, M

文献摘要

被引文献

相似文献

每当系统发育方法应用于潜在的重组核苷酸序列时,对重组程度的准确估计是重要的。在这里,使用一种新的方法检测和测量重组,检查来自病毒,细菌和线粒体的数据集与克隆性的偏差。序列比对中位点间的表观速率异质性(ARH)可以被夸大为重组的假象。然而,多态性位点的组成将不同的数据集与重组产生的ARH与克隆数据集,表现出同等程度的速率异质性。这是因为重组数据集,包括冲突的系统发育历史的区域,往往会产生“星形”的树,表面上类似于从克隆数据集推断的弱系统发育信号。具体而言,与支持相同树形的克隆生成的数据集相比,重组数据集将意外地富含冲突的系统发育信息。它的q值定义为双态简约信息位点占所有多态位点的比例将大于非重组数据的预期值。这里提出的方法,信息网站测试,比较q值对零分布的值,发现使用Monte Carlo模拟的数据进化下的克隆性的零假设。q的显著过量表明克隆性的假设是无效的,因此数据中的ARH至少部分是重组的假象。使用模拟序列对该程序的研究表明,它可以成功地检测和测量重组,并且不太可能产生“假阳性”。模拟还表明,对于重组数据,天真地使用包含速率异质性的最大似然模型可能会导致高估最近共同祖先的时间。将该检验应用于真实的数据,首次揭示了病毒种群与细菌种群一样,可以通过普遍重组接近完全连锁平衡。另一方面,该测试并没有拒绝克隆性的假设时,适用于从人类线粒体DNA的编码区的数据集,尽管其高水平的ARH和同源性。
An accurate estimate of the extent of recombination is important whenever phylogenetic methods are applied to potentially recombining nucleotide sequences. Here, data sets from viruses, bacteria, and mitochondria were examined for deviations from clonality using a new approach for detecting and measuring recombination. The apparent rate heterogeneity (ARH) among sites in a sequence alignment can be inflated as an artifact of recombination. However, the composition of polymorphic sites will differ in a data set with recombination-generated ARH versus a clonal data set that exhibits the equivalent degree of rate heterogeneity. This is because recombinant data sets, encompassing regions of conflicting phylogenetic history, tend to yield "starlike" trees that are superficially similar to those inferred from clonal data sets with weak phylogenetic signal throughout. Specifically, a recombinant data set will be unexpectedly rich in conflicting phylogenetic information compared with clonally generated data sets supporting the same tree shape. Its value of q-defined as the proportion of two-state parsimony-informative sites to all polymorphic sites-will be greater than that expected for nonrecombinant data. The method proposed here, the informative-sites test, compares the value of q against a null distribution of values found using Monte Carlo-simulated data evolved under the null hypothesis of clonality. A significant excess of q indicates that the assumption of clonality is not valid and hence that the ARH in the data is at least partly an artifact of recombination. Investigations of the procedure using simulated sequences indicated that it can successfully detect and measure recombination and that it is unlikely to produce "false positives." Simulations also showed that for recombinant data, naive use of maximum-likelihood models incorporating rate heterogeneity can lead to overestimation of the time to the most recent common ancestor. Application of the test to real data revealed for the first time that populations of viruses, like those of bacteria, can be brought close to complete linkage equilibrium by pervasive recombination. On the other hand, the test did not reject the hypothesis of clonality when applied to a data set from the coding region of human mitochondrial DNA, despite its high level of ARH and homoplasy.