Selecting Question-Specific Genes to Reduce Incongruence in Phylogenomics: A Case Study of Jawed Vertebrate Backbone Phylogeny

Selecting Question-Specific Genes to Reduce Incongruence in Phylogenomics: A Case Study of Jawed Vertebrate Backbone Phylogeny
复制标题

选择特定于问题的基因以减少系统基因组学中的不一致:有颌脊椎动物骨干系统发育的案例研究

DOI:
10.1093/sysbio/syv059
复制
发表时间:
2015-11-01
期刊:
影响因子:
6.5
通讯作者:
Zhang, Peng
Zhang, Peng
中科院分区:
生物学1区
文献类型:
--
作者:
Chen, Meng-Yun;Liang, Dan;Zhang, Peng

文献摘要

被引文献

相似文献

在基因组时代,不同的基因组分析之间的不一致性是遗传学家面临的主要挑战。为了减少不一致性,基因组学研究通常采用一些数据过滤方法,如减少缺失数据或使用缓慢进化的基因,以提高数据的信号质量。在这里,我们组装了一个包含58个有颌脊椎动物分类群和4682个基因的脊椎动物基因组数据集,以研究有颌脊椎动物在连接和合并框架下的骨架遗传。为了评估不同数据过滤方法提取系统发育信号的效率,我们选择了有颌脊椎动物主干系统发育中的六个高度棘手的节间作为我们的测试问题。我们发现,我们的基因组数据集在这些问题的基因之间表现出实质性的冲突信号。我们的分析表明,当一个基因组中有几个困难的节点时,生成的非特定数据集不足以产生一致的结果。此外,基于非特异性数据的系统发育准确性受到数据大小和树推断方法选择的影响。为了解决这样的不一致,我们选择了基因,解决一个给定的节间,但不是整个生育期。值得注意的是,这种策略不仅可以为问题产生正确的关系,而且还可以减少与数据大小和推理方法相关的不一致性。我们的研究强调了基因选择在基因组分析中的重要性,这表明简单地使用大量数据并不能保证正确的结果。构建特定于问题的数据集可能更有助于解决有问题的节点。
Incongruence between different phylogenomic analyses is the main challenge faced by phylogeneticists in the genomic era. To reduce incongruence, phylogenomic studies normally adopt some data filtering approaches, such as reducing missing data or using slowly evolving genes, to improve the signal quality of data. Here, we assembled a phylogenomic data set of 58 jawed vertebrate taxa and 4682 genes to investigate the backbone phylogeny of jawed vertebrates under both concatenation and coalescent-based frameworks. To evaluate the efficiency of extracting phylogenetic signals among different data filtering methods, we chose six highly intractable internodes within the backbone phylogeny of jawed vertebrates as our test questions. We found that our phylogenomic data set exhibits substantial conflicting signal among genes for these questions. Our analyses showed that non-specific data sets that are generated without bias toward specific questions are not sufficient to produce consistent results when there are several difficult nodes within a phylogeny. Moreover, phylogenetic accuracy based on non-specific data is considerably influenced by the size of data and the choice of tree inference methods. To address such incongruences, we selected genes that resolve a given internode but not the entire phylogeny. Notably, not only can this strategy yield correct relationships for the question, but it also reduces inconsistency associated with data sizes and inference methods. Our study highlights the importance of gene selection in phylogenomic analyses, suggesting that simply using a large amount of data cannot guarantee correct results. Constructing question-specific data sets may be more powerful for resolving problematic nodes.