Evaluation of methods for detecting recombination from DNA sequences: Computer simulations

Evaluation of methods for detecting recombination from DNA sequences: Computer simulations
复制标题

DOI:
10.1073/pnas.241370698
复制
发表时间:
2001-11-20
影响因子:
11.1
通讯作者:
Crandall, KA
Crandall, KA
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Posada, D;Crandall, KA

文献摘要

被引文献

相似文献

进化是一个关键的进化过程,它塑造了基因组的结构和种群的遗传结构。虽然许多统计方法可用于检测DNA序列的重组,但它们的绝对和相对性能仍然是未知的。在这里,我们评估了14种不同的重组检测算法的性能。我们使用合并与重组来模拟不同水平的重组,遗传多样性,和网站之间的变化率的DNA序列。将重组检测方法应用于这些数据集,并记录它们是否检测到重组。不同的重组方法表现出不同的性能取决于重组量,遗传多样性,和网站之间的变化率。生成数据的核苷酸取代模型似乎没有显著影响。大多数方法增加功率与更多的序列分歧。一般来说,重组检测方法似乎可以捕获重组的存在,但它们不是很强大。使用替代模式或位点之间的不相容性的方法比基于系统发育不一致的方法更强大。大多数方法似乎不会偶然推断出比预期更多的假阳性。特别是根据数据的多样性,可以使用不同的方法来获得最大的功率,同时最大限度地减少误报。此处显示的结果将为选择最合适的方法分析手头的特定数据提供一些指导。
Recombination is a key evolutionary process that shapes the architecture of genomes and the genetic structure of populations. Although many statistical methods are available for the detection of recombination from DNA sequences, their absolute and relative performance is still unknown. Here we evaluated the performance of 14 different recombination detection algorithms. We used the coalescent with recombination to simulate DNA sequences with different levels of recombination, genetic diversity, and rate variation among sites. Recombination detection methods were applied to these data sets, and whether they detected or not recombination was recorded. Different recombination methods showed distinct performance depending on the amount of recombination, genetic diversity, and rate variation among sites. The model of nucleotide substitution under which the data were generated did not seem to have a significant effect. Most methods increase power with more sequence divergence. In general, recombination detection methods seem to capture the presence of recombination, but they are not very powerful. Methods that use substitution patterns or incompatibility among sites were more powerful than methods based on phylogenetic incongruence. Most methods do not seem to infer more false positives than expected by chance. Especially depending on the amount of diversity in the data, different methods could be used to attain maximum power while minimizing false positives. Results shown here will provide some guidance in the selection of the most appropriate method/s for the analysis of the particular data at hand.