Estimating Gene Gain and Loss Rates in the Presence of Error in Genome Assembly and Annotation Using CAFE 3

Estimating Gene Gain and Loss Rates in the Presence of Error in Genome Assembly and Annotation Using CAFE 3
复制标题

DOI:
10.1093/molbev/mst100
复制
发表时间:
2013-08-01
影响因子:
10.7
通讯作者:
Hahn, Matthew W.
Hahn, Matthew W.
中科院分区:
生物学1区
文献类型:
--
作者:
Han, Mira V.;Thomas, Gregg W. C.;Hahn, Matthew W.

文献摘要

被引文献

相似文献

目前的测序方法产生大量数据,但从这些数据构建的基因组组装通常是片段化和不完整的。不完整和充满错误的组装会导致许多注释错误,特别是在基因组中存在的基因数量方面。这意味着试图估计基因复制和丢失率的方法经常会被这些错误所误导,并且基因家族进化的速率将被一贯高估。在这里,我们提出了一种方法,考虑到这些错误,允许人们准确地推断基因组之间的基因获得和丢失率,即使低组装和注释质量。该方法是在最新版本的软件包CAFE,沿着与其他几个新的功能。我们证明了广泛的模拟和重新分析几个以前发表的数据集的方法的准确性。我们的研究结果表明,基因组注释中的错误确实会导致更高的基因获得和丢失的推断率,但CAFE 3足以解释这些错误,以提供重要的进化参数的准确估计。
Current sequencing methods produce large amounts of data, but genome assemblies constructed from these data are often fragmented and incomplete. Incomplete and error-filled assemblies result in many annotation errors, especially in the number of genes present in a genome. This means that methods attempting to estimate rates of gene duplication and loss often will be misled by such errors and that rates of gene family evolution will be consistently overestimated. Here, we present a method that takes these errors into account, allowing one to accurately infer rates of gene gain and loss among genomes even with low assembly and annotation quality. The method is implemented in the newest version of the software package CAFE, along with several other novel features. We demonstrate the accuracy of the method with extensive simulations and reanalyze several previously published data sets. Our results show that errors in genome annotation do lead to higher inferred rates of gene gain and loss but that CAFE 3 sufficiently accounts for these errors to provide accurate estimates of important evolutionary parameters.