Further comments on the statistical analysis of DNA-DNA hybridization data.

Further comments on the statistical analysis of DNA-DNA hybridization data.
复制标题

对DNA-DNA杂交数据统计分析的进一步评论。

DOI:
10.1093/oxfordjournals.molbev.a040397
复制
发表时间:
1986
影响因子:
10.7
通讯作者:
Templeton,AR
Templeton,AR
中科院分区:
生物学1区
文献类型:
--
作者:
Templeton,AR

文献摘要

被引文献

相似文献

saiitou(1986)和Ruvolo and Smith(1986)对我提出的delta q检验(Templeton 1985)提出了批评,该检验用于测试DNA-DNA杂交数据区分两种不同系统发生的能力。这些批评主要分为两大类:(1)关于δ q检验的统计性质的批评;(2)声称DNA-DNA杂交数据优于其他类型的分子数据的批评。Saitou的批评涉及到δ q检验的统计特性,他的第一个批评是零分布不包括遗传距离数据中常见的层次结构。这种批评再次出现在他关于添加一个假设的黑猩猩物种的讨论中。这个假设的例子使用了这样一个事实,即q检验的幂取决于样本中包含的信息分类群的数量。就其本身而言,权力依赖于信息分类群的数量是该测试的一个合理和理想的性质,但Saitou指出,由于信息分类群之间的等级关系,权力可能会以特殊的方式受到影响。幸运的是,很容易将层次结构合并到δ q检验的零分布中,以消除这些困难。在delta Q测试论文的原始草稿中,我以两种不同的方式推导了delta Q的概率分布——一种(最终发表的)没有分层结构进入零分布,另一种通过随机排列分类群(而不是单个距离条目)来完成推导。Pielou(1979)提出了零分布的两种不同定义,根据她的定义,我将它们分别称为“主要随机性”和“次要随机性”。正如Pielou(1979)所强调的,二次随机s保留了原始距离矩阵中存在的层次结构。不幸的是,关于二次随机性的讨论从已发表的版本中删除了,因为关于Sibley和Ahlquist(1984)的数据得出的结论-即不能拒绝没有歧视的原假设-在任意一种随机性定义下都是相同的。因此,审稿人认为这篇论文没有增加什么新内容,应该局限于更简单的随机性定义。因此,我欢迎Saitou的批评,因为它为重新引入Pielou的δ q统计量的二次随机性概念提供了机会。Saitou的下一个批评是,δ q检验不足以测试系统发生A与系统发生B(使用与Saitou图1中相同的标签),因为除非d43小于dd1或zyxwvutsrqponmlkjihgfedcbaZYX d42,否则人们无法接受5%水平的系统发生A。Saitou认为,对这个不等式的依赖使得q检验不充分,因为即使系统发育a是真的,这个不等式也有“高概率”发生。然而,人们应该记住,q检验的目的不是为了证实系统发生A是真实的,而是为了看看距离数据是否可以区分系统发生A和B,这两个系统发生A和B都不是先验的。在两种系统发生下具有不同概率的任何事件都具有信息性
Saitou (1986) and Ruvolo and Smith (1986) have criticized the delta Q-test that I proposed (Templeton 1985) for testing the ability of DNA-DNA hybridization data to discriminate between two alternative phylogenies. The criticisms fall into two basic categories:(1) those concerning the statistical properties of the delta Q-test and (2) those claiming that DNA-DNA hybridization data are superior to alternative types of molecular data.Saitou’s critique concerns the statistical properties of the delta Q-test, and his first criticism is that the null distribution does not include the hierarchical structure that is commonly found in genetic-distance data. This criticism reappears in his discussion of adding a hypothetical chimpanzee species. This hypothetical example uses the fact that the power of the delta Q-test depends on the number of informative taxa contained in the sample. By itself, the dependency of power on the number of informative taxa is a reasonable and desirable property of the test, but Saitou points out that the power can be affected in peculiar ways because of the hierarchical relationships among the informative taxa. Fortunately, it is very easy to incorporate hierarchical structure into the null distribution of the delta Q-test to eliminate these difficulties. In the original draft of the delta Q-test paper, I derived the probability distribution o f delta Q in two different fashions-one (that which was ultimately published) in which no hierarchical structure enters into the null distribution and another that accomplish es the derivation by randomly permuting the taxa (not individual distance entries). The se two different definitions of the null distribution were suggested by Pielou (1979) and, following her definitions, I referred to them respectively as “primary randomness” and “secondary randomness.” As emphasized by Pielou (1979), secondary randomne ss preserves the hierarchical structure present in the original distance matrix. Unfortunately, the discussion of secondary randomness was deleted from the published versi on because the conclusion-namely, that the null hypothesis imputing no discrimination could not be rejected-reached concerning the Sibley and Ahlquist (1984) data was the same under either definition of randomness. Hence, the reviewers felt that nothing new was added and that the paper should be restricted to the simpler definition of randomness. I therefore welcome Saitou’s criticism since it affords an opportunity to reintroduce Pielou’s concept of secondary randomness for the delta Q-statistic. Saitou’s next criticism is that the delta Q-test is inadequate for testing phylogeny A versus phylogeny B (using the same labels as in Saitou’s fig. 1) because one cannot accept phylogeny A at the 5% level unless d43 is less than ddl or zyxwvutsrqponmlkjihgfedcbaZYX d42. Saitou feels that the dependency on this inequality makes the delta Q-test inadequate because this inequality has a “high probability” of occurring even if phylogeny A is true. However, one should keep in mind that the purpose of the delta Q-test is not to provide confirmation that phylogeny A is true but rather to see whether the distance data can discriminate between phylogenies A and B, neither of which is known to be true a priori. Any event that has different probabilities under the two phylogenies is informative