Human contamination in bacterial genomes has created thousands of spurious proteins

Human contamination in bacterial genomes has created thousands of spurious proteins
复制标题

DOI:
10.1101/gr.245373.118
复制
发表时间:
2019-06-01
期刊:
影响因子:
7
通讯作者:
Salzberg, Steven L.
Salzberg, Steven L.
中科院分区:
生物学1区
文献类型:
--
作者:
Breitwieser, Florian P.;Pertea, Mihaela;Salzberg, Steven L.

文献摘要

被引文献

相似文献

在已发表的基因组中出现的污染物序列可能会给下游分析,特别是进化研究和宏基因组学项目带来许多问题。我们对NCBI RefSeq数据库中完整的和草图的细菌和古细菌基因组进行了大规模扫描,发现2250个基因组受到了人类序列的污染。污染物序列主要来自高拷贝的人类重复序列区域,这些区域本身在当前的人类参考基因组GRCh38中没有充分的代表。人类组装序列的缺失为它们在细菌组装中存在提供了一个可能的解释。在某些情况下,污染的contigs被错误地注释为含有蛋白质编码序列,随着时间的推移,这些序列已经繁殖,在多个原核和真核基因组中产生虚假的蛋白质“家族”。因此,目前在广泛使用的nr和TrEMBL蛋白数据库中存在3437个虚假蛋白条目。我们在这里报告了细菌基因组组装中的污染物序列和与它们相关的蛋白质的广泛列表。我们发现几乎所有的污染物都发生在草图基因组中的小contigs中,这表明从草图基因组组装中过滤掉小contigs可能会减轻污染问题,同时仍然保留几乎所有的真实基因组序列。
Contaminant sequences that appear in published genomes can cause numerous problems for downstream analyses, particularly for evolutionary studies and metagenomics projects. Our large-scale scan of complete and draft bacterial and archaeal genomes in the NCBI RefSeq database reveals that 2250 genomes are contaminated by human sequence. The contaminant sequences derive primarily from high-copy human repeat regions, which themselves are not adequately represented in the current human reference genome, GRCh38. The absence of the sequences from the human assembly offers a likely explanation for their presence in bacterial assemblies. In some cases, the contaminating contigs have been erroneously annotated as containing protein-coding sequences, which over time have propagated to create spurious protein "families" across multiple prokaryotic and eukaryotic genomes. As a result, 3437 spurious protein entries are currently present in the widely used nr and TrEMBL protein databases. We report here an extensive list of contaminant sequences in bacterial genome assemblies and the proteins associated with them. We found that nearly all contaminants occurred in small contigs in draft genomes, which suggests that filtering out small contigs from draft genome assemblies may mitigate the issue of contamination while still keeping nearly all of the genuine genomic sequences.