Introducing SPeDE: High-Throughput Dereplication and Accurate Determination of Microbial Diversity from Matrix-Assisted Laser Desorption-Ionization Time of Flight Mass Spectrometry Data

Introducing SPeDE: High-Throughput Dereplication and Accurate Determination of Microbial Diversity from Matrix-Assisted Laser Desorption-Ionization Time of Flight Mass Spectrometry Data
复制标题

DOI:
10.1128/msystems.00437-19
复制
发表时间:
2019-09-01
期刊:
影响因子:
6.4
通讯作者:
Carlier, Aurelien
Carlier, Aurelien
中科院分区:
生物学2区
文献类型:
--
作者:
Dumolin, Charles;Aerts, Maarten;Carlier, Aurelien

文献摘要

被引文献

相似文献

从微生物群落样品中分离微生物通常会产生大量同种分离株。增加分离株集合所涵盖的多样性需要实施方法和方案以最大限度地减少冗余分离株的数量。基质辅助激光解吸电离飞行时间 (MALDI-TOF) 质谱方法因其低成本和高通量而非常适合解决这种去重复问题。然而,可用的软件工具很麻烦,并且依赖于先前开发的参考数据库或全局相似性分析,这不方便且分类分辨率低。我们推出 SPeDE,这是一种用户友好的光谱数据分析工具,用于 MALDI-TOF 质谱的去重复。 SPeDE 不是依靠全局相似性方法对光谱进行分类,而是通过全局和局部峰值比较的组合来确定独特光谱特征的数量。这种方法允许识别一组与操作隔离单元相关的非冗余光谱。我们在代表 6 个门 132 个属 167 个细菌菌株的 5,228 个光谱的数据集以及冻干和传代培养前后测量的 78 个菌株的 312 个光谱的数据集上评估了 SPeDE。 SPeDE 能够通过识别冗余光谱同时检索样品中所有菌株的参考光谱来高效去除重复。 SPeDE可以识别光谱之间的区别特征,其性能在速度和精度上都超过了现有方法。 SPeDE 是根据 MIT 许可开源的,可从 https://github.com/LM-UGent/SPeDE 获取。重要性 MALDI-TOF 质谱数据集中存在的操作隔离单元的估计涉及一个重要的去重复步骤,以快速识别冗余光谱,而不牺牲生物分辨率。我们描述了 SPeDE,这是一种促进依赖于培养的临床或环境研究的新算法。 SPeDE 能够对分离株进行快速分析和去复制,当培养物的长期储存有限或不可行时,这是一个关键功能。我们证明 SPeDE 可以在物种或菌株水平上有效地识别相似光谱集,超过了其他方法的分类分辨率。与传统的基于基因标记的测序方法相比,MALDI-TOF 质谱和 SPeDE 去重复的高通量容量、速度和低成本应有助于采用培养组学方法进行细菌分离活动。
The isolation of microorganisms from microbial community samples often yields a large number of conspecific isolates. Increasing the diversity covered by an isolate collection entails the implementation of methods and protocols to minimize the number of redundant isolates. Matrix-assisted laser desorption-ionization time-of-flight (MALDI-TOF) mass spectrometry methods are ideally suited to this dereplication problem because of their low cost and high throughput. However, the available software tools are cumbersome and rely either on the prior development of reference databases or on global similarity analyses, which are inconvenient and offer low taxonomic resolution. We introduce SPeDE, a user-friendly spectral data analysis tool for the dereplication of MALDI-TOF mass spectra. Rather than relying on global similarity approaches to classify spectra, SPeDE determines the number of unique spectral features by a mix of global and local peak comparisons. This approach allows the identification of a set of nonredundant spectra linked to operational isolation units. We evaluated SPeDE on a data set of 5,228 spectra representing 167 bacterial strains belonging to 132 genera across six phyla and on a data set of 312 spectra of 78 strains measured before and after lyophilization and subculturing. SPeDE was able to dereplicate with high efficiency by identifying redundant spectra while retrieving reference spectra for all strains in a sample. SPeDE can identify distinguishing features between spectra, and its performance exceeds that of established methods in speed and precision. SPeDE is open source under the MIT license and is available from https://github.com/LM-UGent/SPeDE.IMPORTANCE Estimation of the operational isolation units present in a MALDI-TOF mass spectral data set involves an essential dereplication step to identify redundant spectra in a rapid manner and without sacrificing biological resolution. We describe SPeDE, a new algorithm which facilitates culture-dependent clinical or environmental studies. SPeDE enables the rapid analysis and dereplication of isolates, a critical feature when long-term storage of cultures is limited or not feasible. We show that SPeDE can efficiently identify sets of similar spectra at the level of the species or strain, exceeding the taxonomic resolution of other methods. The high-throughput capacity, speed, and low cost of MALDI-TOF mass spectrometry and SPeDE dereplication over traditional gene marker-based sequencing approaches should facilitate adoption of the culturomics approach to bacterial isolation campaigns.