Machine learning classification by fitting amplicon sequences to existing OTUs.

Machine learning classification by fitting amplicon sequences to existing OTUs.
复制标题

DOI:
10.1128/msphere.00336-23
复制
发表时间:
2023-10-24
期刊:
影响因子:
4.8
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

使用16S rRNA基因序列数据来训练机器学习分类模型的能力提供了根据微生物组组成诊断患者的机会。在某些应用中,提供最佳模型的分类解析可能需要使用从头运算分类单位(OTU),其组成在添加新数据时发生变化。我们之前开发了一种新的基于参考的方法OptiFit,该方法将新的序列数据拟合到现有的从头OTU,而不改变原始OTU的组成。虽然OptiFit产生的OTU与de novo OTU一样高质量,但尚不清楚这种将新序列数据拟合到现有OTU中的方法是否会影响分类模型相对于仅使用de novo OTU训练和测试的模型的性能。我们使用OptiFit将序列聚类到现有的OTU中,并评估模型在对包含来自患有和不患有结肠筛查相关瘤形成(SRN)的患者的样本的数据集进行分类时的性能。我们将该模型的性能与标准方法进行了比较,包括从头和基于数据库引用的聚类。我们发现,使用OptiFit在分类SRN方面表现良好或更好。OptiFit可以通过避免使用重新聚类的序列重新训练模型来简化对新样本进行分类的过程。使用微生物组数据来帮助诊断有很大的潜力。基于从头操作分类单元(OTU)的分类模型的挑战在于,16S rRNA基因序列通常基于与数据集中的其他序列的相似性而被分配给OTU。如果从新患者生成数据,则必须将旧序列和新序列重新聚类到OTU,并重新训练分类模型。然而,人们希望有一个单一的,经过验证的模型,可以广泛部署。为了克服这一障碍,我们应用OptiFit聚类算法将新的序列数据拟合到现有的OTU,从而允许重复使用模型。使用OptiFit实现的随机森林模型与传统的重新分配和重新训练方法一样好。该结果表明,可以基于OTU相对丰度数据训练和应用机器学习模型,而不需要重新训练或使用参考数据库。
The ability to use 16S rRNA gene sequence data to train machine learning classification models offers the opportunity to diagnose patients based on the composition of their microbiome. In some applications, the taxonomic resolution that provides the best models may require the use of de novo operational taxonomic units (OTUs) whose composition changes when new data are added. We previously developed a new reference-based approach, OptiFit, that fits new sequence data to existing de novo OTUs without changing the composition of the original OTUs. While OptiFit produces OTUs that are as high quality as de novo OTUs, it is unclear whether this method for fitting new sequence data into existing OTUs will impact the performance of classification models relative to models trained and tested only using de novo OTUs. We used OptiFit to cluster sequences into existing OTUs and evaluated model performance in classifying a dataset containing samples from patients with and without colonic screen relevant neoplasia (SRN). We compared the performance of this model to standard methods including de novo and database-reference-based clustering. We found that using OptiFit performed as well or better in classifying SRNs. OptiFit can streamline the process of classifying new samples by avoiding the need to retrain models using reclustered sequences. There is great potential for using microbiome data to aid in diagnosis. A challenge with de novo operational taxonomic unit (OTU)-based classification models is that 16S rRNA gene sequences are often assigned to OTUs based on similarity to other sequences in the dataset. If data are generated from new patients, the old and new sequences must be reclustered to OTUs and the classification model retrained. Yet there is a desire to have a single, validated model that can be widely deployed. To overcome this obstacle, we applied the OptiFit clustering algorithm to fit new sequence data to existing OTUs allowing for reuse of the model. A random forest model implemented using OptiFit performed as well as the traditional reassign and retrain approach. This result shows that it is possible to train and apply machine learning models based on OTU relative abundance data that do not require retraining or the use of a reference database.
DOI: 10.1038/s41467-017-01973-8
发表时间: 2017-12-05
影响因子: 16.6
作者:
Duvallet C;Gibbons SM;Gurry T;Irizarry RA;Alm EJ
通讯作者: Alm EJ
DOI: 10.21105/joss.03073
发表时间: 2021-01-01
影响因子: --
作者:
Topcuoglu, Begum D;Lapp, Zena;Schloss, Patrick D
通讯作者: Schloss, Patrick D
DOI: 10.1128/aem.01541-09
发表时间: 2009-12-01
影响因子: 4.4
作者:
Schloss, Patrick D.;Westcott, Sarah L.;Weber, Carolyn F.
通讯作者: Weber, Carolyn F.
DOI: 10.1016/j.chom.2014.02.005
发表时间: 2014-03-12
影响因子: 30.3
作者:
Gevers D;Kugathasan S;Denson LA;Vázquez-Baeza Y;Van Treuren W;Ren B;Schwager E;Knights D;Song SJ;Yassour M;Morgan XC;Kostic AD;Luo C;González A;McDonald D;Haberman Y;Walters T;Baker S;Rosh J;Stephens M;Heyman M;Markowitz J;Baldassano R;Griffiths A;Sylvester F;Mack D;Kim S;Crandall W;Hyams J;Huttenhower C;Knight R;Xavier RJ
通讯作者: Xavier RJ
DOI: 10.1186/s13073-016-0290-3
发表时间: 2016-04-06
期刊: Genome medicine
影响因子: 12.3
作者:
Baxter NT;Ruffin MT 4th;Rogers MA;Schloss PD
通讯作者: Schloss PD