The impact of transitive annotation on the training of taxonomic classifiers.

The impact of transitive annotation on the training of taxonomic classifiers.
复制标题

DOI:
10.3389/fmicb.2023.1240957
复制
发表时间:
2023
影响因子:
5.2
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

微生物群落分析中的一项常见任务涉及将分类标签分配给来自群落中发现的生物体的序列。通常,使用机器学习算法来分配这样的标签,所述机器学习算法被训练以基于包括具有已知分类标签的序列的训练数据集来识别各个分类组。理想情况下,训练数据应该依赖于实验验证的标签-正式的分类标签需要生物体的物理和生化特性的知识,这些知识不能仅从序列直接推断。然而,与生物数据库中的序列相关联的标签是最常见的计算预测,其本身可能依赖于计算生成的数据-一个通常被称为“传递注释”的过程。在这篇手稿中,我们探讨了在计算生成的数据上训练机器学习分类器(在我们的案例中是核糖体数据库项目的贝叶斯分类器)的影响。我们基于来自宏基因组实验的16 S rRNA数据生成新的训练样本,并评估分类器预测的分类标签在重新训练后的变化程度。我们证明,即使是几个计算生成的训练数据点可以显着倾斜的分类器的输出到分类空间的整个区域可以被干扰的点。最后,我们讨论了影响分类器对transitively-annotated训练数据的弹性的关键因素,并提出了避免本文中描述的伪影的最佳实践。
A common task in the analysis of microbial communities involves assigning taxonomic labels to the sequences derived from organisms found in the communities. Frequently, such labels are assigned using machine learning algorithms that are trained to recognize individual taxonomic groups based on training data sets that comprise sequences with known taxonomic labels. Ideally, the training data should rely on labels that are experimentally verified—formal taxonomic labels require knowledge of physical and biochemical properties of organisms that cannot be directly inferred from sequence alone. However, the labels associated with sequences in biological databases are most commonly computational predictions which themselves may rely on computationally-generated data—a process commonly referred to as “transitive annotation.” In this manuscript we explore the implications of training a machine learning classifier (the Ribosomal Database Project’s Bayesian classifier in our case) on data that itself has been computationally generated. We generate new training examples based on 16S rRNA data from a metagenomic experiment, and evaluate the extent to which the taxonomic labels predicted by the classifier change after re-training. We demonstrate that even a few computationally-generated training data points can significantly skew the output of the classifier to the point where entire regions of the taxonomic space can be disturbed. We conclude with a discussion of key factors that affect the resilience of classifiers to transitively-annotated training data, and propose best practices to avoid the artifacts described in our paper.
基因组再保管:Wiki解决方案?
DOI: 10.1186/gb-2007-8-1-102
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
通讯作者: --
DOI: 10.1186/s40793-015-0101-2
发表时间: 2015
影响因子: --
作者:
Promponas VJ;Iliopoulos I;Ouzounis CA
通讯作者: Ouzounis CA
DOI: 10.1038/s41396-021-00941-x
发表时间: 2021-07
期刊: The ISME journal
影响因子: --
作者:
Hugenholtz P;Chuvochina M;Oren A;Parks DH;Soo RM
通讯作者: Soo RM
DOI: 10.1128/aem.00062-07
发表时间: 2007-08-01
影响因子: 4.4
作者:
Wang, Qiong;Garrity, George M.;Cole, James R.
通讯作者: Cole, James R.
DOI: 10.1099/ijs.0.02611-0
发表时间: 2003-09-01
影响因子: 2.8
作者:
Li, WJ;Xu, P;Jiang, CL
通讯作者: Jiang, CL