"Reverse ecology" and the power of population genomics.

"Reverse ecology" and the power of population genomics.
复制标题

DOI:
10.1111/j.1558-5646.2008.00486.x
复制
发表时间:
2008-12
期刊:
Evolution; international journal of organic evolution
影响因子:
--
通讯作者:
Hahn MW
Hahn MW
中科院分区:
其他
文献类型:
--
作者:
Li YF;Costello JC;Holloway AK;Hahn MW

文献摘要

参考文献

被引文献

相似文献

快速和廉价的测序技术使得从一个群体中收集多个个体的全基因组序列数据成为可能。这种类型的数据可以通过寻找适应性自然选择的目标来快速识别控制重要生态和进化表型的基因,因此我们将这种方法称为“逆向生态学”。为了量化利用群体基因组数据检测阳性选择所获得的能力,我们比较了三种识别选择目标的统计方法:McDonald-Kreitman检验、mkprf方法和检测dN/dS bbb1的可能性实现。因为前两种方法使用多态性数据,我们期望它们在检测选择方面有更大的能力。然而,当应用于人类、苍蝇和酵母的群体基因组数据集时,使用多态性数据的测试实际上在三个数据集中的两个中较弱。我们探讨了为什么更简单的比较方法在选择中发现了更多的基因,并提出不同的方法可能真的是从相同的序列数据中检测到不同的信号。最后,我们发现了与mkprf方法相关的几个统计异常,包括鉴定的正选择基因数量与使用的先验分布之间几乎呈线性依赖关系。我们的结论是,解释这种方法产生的结果应该谨慎一些。
Rapid and inexpensive sequencing technologies are making it possible to collect whole genome sequence data on multiple individuals from a population. This type of data can be used to quickly identify genes that control important ecological and evolutionary phenotypes by finding the targets of adaptive natural selection, and we therefore refer to such approaches as “reverse ecology.” In order to quantify the power gained in detecting positive selection using population genomic data, we compare three statistical methods for identifying targets of selection: the McDonald-Kreitman test, the mkprf method, and a likelihood implementation for detecting dN/dS>1. Because the first two methods use polymorphism data we expect them to have more power to detect selection. However, when applied to population genomic datasets from human, fly, and yeast, the tests using polymorphism data were actually weaker in two of the three datasets. We explore reasons why the simpler comparative method has identified more genes under selection, and suggest that the different methods may really be detecting different signals from the same sequence data. Finally, we find several statistical anomalies associated with the mkprf method, including an almost linear dependence between the number of positively selected genes identified and the prior distributions used. We conclude that interpreting the results produced by this method should be done with some caution.
DOI: 10.1073/pnas.0409159102
发表时间: 2005-01-25
影响因子: 11.1
作者:
Gu, ZL;David, L;Steinmetz, LM
通讯作者: Steinmetz, LM
DOI: 10.1371/journal.pbio.0020286
发表时间: 2004-10
期刊: PLoS biology
影响因子: 9.8
作者:
Akey JM;Eberle MA;Rieder MJ;Carlson CS;Shriver MD;Nickerson DA;Kruglyak L
通讯作者: Kruglyak L
DOI: 10.1038/416531a
发表时间: 2002-04-04
期刊: NATURE
影响因子: 64.8
作者:
Bustamante, CD;Nielsen, R;Hartl, DL
通讯作者: Hartl, DL
DOI: 10.1038/nature01644
发表时间: 2003-05-15
期刊: NATURE
影响因子: 64.8
作者:
Kellis, M;Patterson, N;Lander, ES
通讯作者: Lander, ES
DOI: 10.1038/4151024a
发表时间: 2002-02-28
期刊: NATURE
影响因子: 64.8
作者:
Fay, JC;Wyckoff, GJ;Wu, CI
通讯作者: Wu, CI