Measuring Phylogenetic Information of Incomplete Sequence Data

Measuring Phylogenetic Information of Incomplete Sequence Data
复制标题

DOI:
10.1093/sysbio/syab073
复制
发表时间:
2021-10-29
期刊:
影响因子:
6.5
通讯作者:
Thorne, Jeffrey L.
Thorne, Jeffrey L.
中科院分区:
生物学1区
文献类型:
--
作者:
Seo, Tae-Kun;Gascuel, Olivier;Thorne, Jeffrey L.

文献摘要

被引文献

相似文献

从比对的分子序列集合中提取系统发育信息的广泛使用的方法依赖于核苷酸替换或氨基酸替换的概率模型。可以提取的系统发育信息取决于序列比对中的列数,并且当比对包含由于插入或缺失事件引起的空位时,系统发育信息将减少。出于信息损失的测量,我们建议评估对齐数据集的有效序列长度(ESL)。由于比对缺口的存在,ESL可以不同于序列比对中的实际列数。此外,系统发育信息的估计受到模型误设的影响。不可避免的是,分子进化的实际过程不同于用来描述这一过程的概率模型。这种差异意味着实际序列比对中的系统发育信息量将不同于相同大小的模拟数据集中的信息量,这促使我们开发一种新的模型充分性测试。通过理论和实证数据分析,我们展示了如何理清差距和模型误设的影响。通过比较真实序列和模拟序列的Fisher信息,我们确定哪些比对位点和分支最受空位和模型错误指定的影响。[Fisher信息;缺口;插入;缺失; indel;模型充分性;拟合优度检验;序列比对。]
Widely used approaches for extracting phylogenetic information from aligned sets of molecular sequences rely upon probabilistic models of nucleotide substitution or amino-acid replacement. The phylogenetic information that can be extracted depends on the number of columns in the sequence alignment and will be decreased when the alignment contains gaps due to insertion or deletion events. Motivated by the measurement of information loss, we suggest assessment of the effective sequence length (ESL) of an aligned data set. The ESL can differ from the actual number of columns in a sequence alignment because of the presence of alignment gaps. Furthermore, the estimation of phylogenetic information is affected by model misspecification. Inevitably, the actual process of molecular evolution differs from the probabilistic models employed to describe this process. This disparity means the amount of phylogenetic information in an actual sequence alignment will differ from the amount in a simulated data set of equal size, which motivated us to develop a new test for model adequacy. Via theory and empirical data analysis, we show how to disentangle the effects of gaps and model misspecification. By comparing the Fisher information of actual and simulated sequences, we identify which alignment sites and tree branches are most affected by gaps and model misspecification. [Fisher information; gaps; insertion; deletion; indel; model adequacy; goodness-of-fit test; sequence alignment.]