Detecting and correcting misclassified sequences in the large-scale public databases.

Detecting and correcting misclassified sequences in the large-scale public databases.
复制标题

DOI:
10.1093/bioinformatics/btaa586
复制
发表时间:
2020-09-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Rajan H
Rajan H
中科院分区:
其他
文献类型:
--
作者:
Bagheri H;Severin AJ;Rajan H

文献摘要

参考文献

被引文献

相似文献

随着测序成本的降低,存放在公共存储库中的数据量正在迅速增加。公共数据库依赖用户为每次提交的数据提供元数据,这很容易导致用户出错。不幸的是,大多数公共数据库,如非冗余(NR),依赖于用户输入,并没有方法来识别所提供的元数据中的错误,导致错误传播的可能性。以前的研究NR数据库的一个小的子集分析基于序列相似性的误分类。据我们所知,整个数据库中的分类错误数量尚未量化。我们提出了一种启发式方法来检测潜在的错误分类NR数据库中的分类分配。我们应用了一种管理技术和质量控制,以找到最可能的分类分配。我们的方法结合了手工和计算创建的数据库和聚类信息在95%的相似性每个注释的出处和频率。我们在NR数据库中发现了200多万种可能分类错误的蛋白质。使用模拟数据,我们表现出97%的高精度和87%的召回检测分类错误的蛋白质。拟议的方法和结论也可适用于其他数据库。源代码、数据集、文档、Docker Notebook和Docker容器可在https://github.com/boalang/nr上获得。 补充数据可在Bioinformatics在线获得。
As the cost of sequencing decreases, the amount of data being deposited into public repositories is increasing rapidly. Public databases rely on the user to provide metadata for each submission that is prone to user error. Unfortunately, most public databases, such as non-redundant (NR), rely on user input and do not have methods for identifying errors in the provided metadata, leading to the potential for error propagation. Previous research on a small subset of the NR database analyzed misclassification based on sequence similarity. To the best of our knowledge, the amount of misclassification in the entire database has not been quantified. We propose a heuristic method to detect potentially misclassified taxonomic assignments in the NR database. We applied a curation technique and quality control to find the most probable taxonomic assignment. Our method incorporates provenance and frequency of each annotation from manually and computationally created databases and clustering information at 95% similarity. We found more than two million potentially taxonomically misclassified proteins in the NR database. Using simulated data, we show a high precision of 97% and a recall of 87% for detecting taxonomically misclassified proteins. The proposed approach and findings could also be applied to other databases. Source code, dataset, documentation, Jupyter notebooks and Docker container are available at https://github.com/boalang/nr. Supplementary data are available at Bioinformatics online.
DOI: 10.7717/peerj.4652
发表时间: 2018
期刊: PeerJ
影响因子: 2.7
作者:
Edgar RC
通讯作者: Edgar RC
DOI: 10.1186/1471-2105-9-353
发表时间: 2008-08-27
期刊: BMC bioinformatics
影响因子: 3
作者:
Nagy A;Hegyi H;Farkas K;Tordai H;Kozma E;Bányai L;Patthy L
通讯作者: Patthy L
DOI: 10.1093/nar/gky359
发表时间: 2018-07-02
影响因子: 14.9
作者:
Medlar AJ;Törönen P;Holm L
通讯作者: Holm L
DOI: 10.1093/bioinformatics/bts565
发表时间: 2012-12-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Fu L;Niu B;Zhu Z;Wu S;Li W
通讯作者: Li W
DOI: 10.1093/nar/gku989
发表时间: 2015-01
影响因子: 14.9
作者:
UniProt Consortium
通讯作者: UniProt Consortium