Tandem repeats lead to sequence assembly errors and impose multi-level challenges for genome and protein databases

Tandem repeats lead to sequence assembly errors and impose multi-level challenges for genome and protein databases
复制标题

DOI:
10.1093/nar/gkz841
复制
发表时间:
2019-12-02
影响因子:
14.9
通讯作者:
Linke, Dirk
Linke, Dirk
中科院分区:
生物学2区
文献类型:
--
作者:
Torresen, Ole K.;Star, Bastiaan;Linke, Dirk

文献摘要

被引文献

相似文献

在整个生命之树的生物体基因组中,DNA重复延伸的广泛存在给测序、基因组组装以及基因和蛋白质的自动注释带来了根本性的挑战。这种多层次的问题可能导致基因组和蛋白质数据库中的错误,而这些错误往往不被识别或承认。因此,使用重复区域序列的最终用户面临着“随时可用”的存储数据,其可信度难以确定,更不用说量化了。在这里,我们回顾了与串联重复序列相关的问题,这些问题起源于测序-组装-注释-沉积工作流程的不同阶段,并且可能在公共数据库存储库中扩散,影响所有下游分析。作为一个案例研究,我们提供了大西洋鳕鱼基因组的例子,其测序和组装受到串联重复序列特别高的流行率的阻碍。我们用其他物种的例子来补充这个案例研究,其中错误注释和测序错误已经传播到蛋白质数据库中。通过这篇综述,我们的目标是提高数据库用户社区的意识水平,并提醒在数据库创建的底层工作流程中工作的科学家,他们遗漏或不正确组装的数据很可能包含对其他人有价值的重要生物信息。
The widespread occurrence of repetitive stretches of DNA in genomes of organisms across the tree of life imposes fundamental challenges for sequencing, genome assembly, and automated annotation of genes and proteins. This multi-level problem can lead to errors in genome and protein databases that are often not recognized or acknowledged. As a consequence, end users working with sequences with repetitive regions are faced with 'ready-to-use' deposited data whose trustworthiness is difficult to determine, let alone to quantify. Here, we provide a review of the problems associated with tandem repeat sequences that originate from different stages during the sequencing-assembly-annotation-deposition workflow, and that may proliferate in public database repositories affecting all downstream analyses. As a case study, we provide examples of the Atlantic cod genome, whose sequencing and assembly were hindered by a particularly high prevalence of tandem repeats. We complement this case study with examples from other species, where mis-annotations and sequencing errors have propagated into protein databases. With this review, we aim to raise the awareness level within the community of database users, and alert scientists working in the underlying workflow of database creation that the data they omit or improperly assemble may well contain important biological information valuable to others.