Detecting natural selection in RNA virus populations using sequence summary statistics

Detecting natural selection in RNA virus populations using sequence summary statistics
复制标题

DOI:
10.1016/j.meegid.2009.06.001
复制
发表时间:
2010-04-01
影响因子:
3.2
通讯作者:
Pybus, Oliver G.
Pybus, Oliver G.
中科院分区:
医学3区
文献类型:
--
作者:
Bhatt, Samir;Katzourakis, Aris;Pybus, Oliver G.

文献摘要

被引文献

相似文献

目前,大多数旨在检测自然选择对病毒基因序列的作用的分析都使用沉默突变与替代突变比率的系统发育估计,然而,这种方法对于包含数百个完整病毒基因组的大数据集进行计算是不切实际的,由于基因组测序技术的进步,这些数据集变得越来越普遍。在这里,我们研究基于序列汇总统计的计算高效测试的统计性能,并以两种方式探索它们对 RNA 病毒数据集的适用性。首先,我们进行广泛的模拟,以测量两种著名的汇总统计方法(Tajima's D 和 McDonald-Kreitman 测试)在一系列病毒样突变和人口统计场景下的 I 型错误。其次,我们将这些方法应用于代表天然 RNA 病毒种群的 100 个 RNA 病毒比对的汇编。此外,我们开发并介绍了 McDonald-Kreitman 检验的新实现,并表明它极大地提高了该检验在典型病毒数据集上的统计可靠性。我们的结果表明,McDonald-Kreitman 检验的变体在分析非常大的高度多样化病毒遗传数据集时可能非常有用 (C) 2009 Elsevier B.V. 保留所有权利。
At present, most analyses that aim to detect the action of natural selection upon viral gene sequences use phylogenetic estimates of the ratio of silent to replacement mutations Such methods, however, are impractical to compute on large data sets comprising hundreds of complete viral genomes, which are becoming increasingly common due to advances in genome sequencing technology. Here we investigate the statistical performance of computationally efficient tests that are based on sequence summary statistics, and explore their applicability to RNA virus data sets in two ways Firstly, we perform extensive simulations in order to measure the type I error of two well-known summary statistic methods - Tajima's D and the McDonald-Kreitman test - under a range of virus-like mutational and demographic scenarios Secondly, we apply these methods to a compilation of 100 RNA virus alignments that represent natural RNA virus populations In addition, we develop and introduce a new implementation of the McDonald-Kreitman test and show that it greatly improves the test's statistical reliability on typical viral data sets Our results suggest that variants of the McDonald-Kreitman test could prove useful in the analysis of very large sets of highly diverse viral genetic data (C) 2009 Elsevier B.V. All rights reserved.