A Monte Carlo Approach Successfully Identifies Randomness in Multiple Sequence Alignments : A More Objective Means of Data Exclusion

A Monte Carlo Approach Successfully Identifies Randomness in Multiple Sequence Alignments : A More Objective Means of Data Exclusion
复制标题

DOI:
10.1093/sysbio/syp006
复制
发表时间:
2009-02-01
期刊:
影响因子:
6.5
通讯作者:
Misof, Katharina
Misof, Katharina
中科院分区:
生物学1区
文献类型:
--
作者:
Misof, Bernhard;Misof, Katharina

文献摘要

被引文献

相似文献

序列或序列部分的随机相似性可能会阻碍系统发育分析或基因同源性的鉴定。此外,随机相似的序列或不明确比对的序列部分可以负面地干扰替代模型参数的估计。系统基因组学研究表明,模型估计和树重建中的偏差即使在大数据集也不会消失。事实上,这些偏见可以随着更多的数据而变得明显。因此,在模型估计和树重建之前识别序列比对中可能的随机相似性是重要的。已经提出了不同的方法来识别和处理有问题的路线部分。我们提出了一种替代方法,可以识别随机相似性多序列比对(MSA)的基础上滑动窗口内的蒙特卡罗响应。该方法从成对序列比较中推断相似性谱,随后计算共有谱。该一致性特征代表所有计算的单一相似性特征的总结。因此,共识配置文件确定主导模式的非随机相似性或随机性内的部分MSA。我们表明,该方法清楚地识别随机模拟和真实的数据。在排除了假定的随机部分之后,节点支持在两个数据的树重建中急剧改善。因此,它似乎是一个强大的工具,以确定可能的偏见树重建或基因鉴定。该方法目前仅限于核苷酸数据,但将在不久的将来扩展到蛋白质数据。
Random similarity of sequences or sequence sections can impede phylogenetic analyses or the identification of gene homologies. Additionally, randomly similar sequences or ambiguously aligned sequence sections can negatively interfere with the estimation of substitution model parameters. Phylogenomic studies have shown that biases in model estimation and tree reconstructions do not disappear even with large data sets. In fact, these biases can become pronounced with more data. It is therefore important to identify possible random similarity within sequence alignments in advance of model estimation and tree reconstructions. Different approaches have been already suggested to identify and treat problematic alignment sections. We propose an alternative method that can identify random similarity within multiple sequence alignments (MSAs) based on Monte Carlo resampling within a sliding window. The method infers similarity profiles from pairwise sequence comparisons and subsequently calculates a consensus profile. This consensus profile represents a summary of all calculated single similarity profiles. In consequence, consensus profiles identify dominating patterns of nonrandom similarity or randomness within sections of MSAs. We show that the approach clearly identifies randomness in simulated and real data. After the exclusion of putative random sections, node support drastically improves in tree reconstructions of both data. It thus appears to be a powerful tool to identify possible biases of tree reconstructions or gene identification. The method is currently restricted to nucleotide data but will be extended to protein data in the near future.