Statistics and Truth in Phylogenomics

Statistics and Truth in Phylogenomics
复制标题

DOI:
10.1093/molbev/msr202
复制
发表时间:
2012-02-01
影响因子:
10.7
通讯作者:
Tamura, Koichiro
Tamura, Koichiro
中科院分区:
生物学1区
文献类型:
--
作者:
Kumar, Sudhir;Filipski, Alan J.;Tamura, Koichiro

文献摘要

被引文献

相似文献

系统基因组学是指使用基因组规模的序列数据推断物种之间的历史关系,并使用系统发育分析来推断多基因家族中的蛋白质功能。随着测序成本的迅速降低,基因组学正成为基因组规模和分类学上密集采样数据集的进化分析的同义词。在系统发育推断应用中,这转化为非常大的数据集,可以产生方差极小和统计置信度(P值)高的进化和功能推断。然而,高度显着的P值的报告越来越多,甚至对比系统发育的假设,这取决于所使用的进化模型和推理方法,很难建立真正的关系。我们认为,评估结果的稳健性的生物因素,可能会系统地误导(偏见)的统计估计的结果,将是一个关键,以避免不正确的基因组推断。事实上,除了零假设统计检验的P值之外,还需要更加强调差异的大小(效应量)。另一方面,可用的序列数据的量对于一些非基因组应用可能总是不足的,例如,涉及在单个密码子位置和特定谱系中的偶发性正选择的那些应用。同样,关注效应量和生物学相关性,而不是P值,可能是有道理的。在这里,我们提出了一个理论概述,并讨论了实际方面的影响大小,偏差和P值之间的相互作用,因为它涉及到的统计推断的进化真理在基因组学。
Phylogenomics refers to the inference of historical relationships among species using genome-scale sequence data and to the use of phylogenetic analysis to infer protein function in multigene families. With rapidly decreasing sequencing costs, phylogenomics is becoming synonymous with evolutionary analysis of genome-scale and taxonomically densely sampled data sets. In phylogenetic inference applications, this translates into very large data sets that yield evolutionary and functional inferences with extremely small variances and high statistical confidence (P value). However, reports of highly significant P values are increasing even for contrasting phylogenetic hypotheses depending on the evolutionary model and inference method used, making it difficult to establish true relationships. We argue that the assessment of the robustness of results to biological factors, that may systematically mislead (bias) the outcomes of statistical estimation, will be a key to avoiding incorrect phylogenomic inferences. In fact, there is a need for increased emphasis on the magnitude of differences (effect sizes) in addition to the P values of the statistical test of the null hypothesis. On the other hand, the amount of sequence data available will likely always remain inadequate for some phylogenomic applications, for example, those involving episodic positive selection at individual codon positions and in specific lineages. Again, a focus on effect size and biological relevance, rather than the P value, may be warranted. Here, we present a theoretical overview and discuss practical aspects of the interplay between effect sizes, bias, and P values as it relates to the statistical inference of evolutionary truth in phylogenomics.