Phylogeny-aware identification and correction of taxonomically mislabeled sequences

Phylogeny-aware identification and correction of taxonomically mislabeled sequences
复制标题

DOI:
10.1093/nar/gkw396
复制
发表时间:
2016-06-20
影响因子:
14.9
通讯作者:
Stamatakis, Alexandros
Stamatakis, Alexandros
中科院分区:
生物学2区
文献类型:
--
作者:
Kozlov, Alexey M.;Zhang, Jiajie;Stamatakis, Alexandros

文献摘要

被引文献

相似文献

公共数据库中的分子序列大多由提交作者注释,无需进一步验证。此过程可能会生成错误的分类序列标签。错误标记的序列很难识别,而且它们可能导致下游错误,因为新的序列通常使用现有的序列进行注释。此外,参考序列数据库中的分类错误标记可能会使依赖于分类的元遗传学研究产生偏差。尽管在提高分类学注释的质量方面做出了很大努力,但由于劳动密集型的人工整理过程,精确率很低。在这里,我们介绍了SAVIA,这是一种系统发育感知的方法,可以使用进化的统计模型自动识别分类错误标记的序列(错误标记)。我们使用进化放置算法(EPA)来检测和评分其分类注释不被潜在的系统发育信号支持的序列,并自动为这些序列提出正确的分类分类。实验结果表明,该方法具有较高的识别准确率(96.9%敏感度/91.7%准确率)和误判正确率(94.9%敏感度/89.9%准确率)。此外,对四个广泛使用的微生物16S参考数据库(Greengenes、LTP、RDP和Silva)的分析表明,它们目前包含0.2%至2.5%的错误标签。最后,我们使用SAVIA对蓝藻的可选分类进行了深入的评估。
Molecular sequences in public databases are mostly annotated by the submitting authors without further validation. This procedure can generate erroneous taxonomic sequence labels. Mislabeled sequences are hard to identify, and they can induce downstream errors because new sequences are typically annotated using existing ones. Furthermore, taxonomic mislabelings in reference sequence databases can bias metagenetic studies which rely on the taxonomy. Despite significant efforts to improve the quality of taxonomic annotations, the curation rate is low because of the labor-intensive manual curation process. Here, we present SATIVA, a phylogeny-aware method to automatically identify taxonomically mislabeled sequences ('mislabels') using statistical models of evolution. We use the Evolutionary Placement Algorithm (EPA) to detect and score sequences whose taxonomic annotation is not supported by the underlying phylogenetic signal, and automatically propose a corrected taxonomic classification for those. Using simulated data, we show that our method attains high accuracy for identification (96.9% sensitivity/91.7% precision) as well as correction (94.9% sensitivity/89.9% precision) of mislabels. Furthermore, an analysis of four widely used microbial 16S reference databases (Greengenes, LTP, RDP and SILVA) indicates that they currently contain between 0.2% and 2.5% mislabels. Finally, we use SATIVA to perform an in-depth evaluation of alternative taxonomies for Cyanobacteria.