A Machine Learning Based Approach to de novo Sequencing of Glycans from Tandem Mass Spectrometry Spectrum

A Machine Learning Based Approach to de novo Sequencing of Glycans from Tandem Mass Spectrometry Spectrum
复制标题

DOI:
10.1109/tcbb.2015.2430317
复制
发表时间:
2015-11-01
影响因子:
4.5
通讯作者:
Sakakibara, Yasubumi
Sakakibara, Yasubumi
中科院分区:
工程技术3区
文献类型:
--
作者:
Kumozaki, Shotaro;Sato, Kengo;Sakakibara, Yasubumi

文献摘要

被引文献

相似文献

近年来,糖组学得到了积极的研究,并且用于糖组学的各种技术得到了迅速发展。目前,串联质谱(MS/MS)是鉴定寡糖结构的关键实验工具之一。MS/MS可以观察到碎片聚糖离子的MS/MS峰,包括内部裂解产生的交叉环离子,这为推断聚糖结构提供了有价值的信息。因此,聚糖从头测序的目的是在没有数据库的情况下找到观察到的MS/MS峰与聚糖亚结构的最可能归属。然而,对于从MS/MS谱进行聚糖从头测序,几乎没有令人满意的算法。我们提出了一种基于机器学习的方法,从MS/MS谱中对聚糖进行从头测序。首先,我们建立了一个合适的模型,包括跨环离子的聚糖片段,并实现了一个求解器,采用拉格朗日松弛与动态规划技术。然后,为了优化算法的得分,我们引入了一种称为结构化支持向量机的机器学习技术,该技术使我们能够从训练数据中学习包括交叉环离子得分的参数,即,已知聚糖质谱。此外,我们还对已知聚糖类型(包括N-连接聚糖和O-连接聚糖)的核心结构实施了额外的约束。如果已知给定光谱的聚糖类型,这使我们能够预测更准确的聚糖结构。计算实验表明,我们的算法进行准确的从头测序的聚糖。我们的算法和数据集的实现可在http://glyfon.dna.bio.keio.ac.jp/上获得。
Recently, glycomics has been actively studied and various technologies for glycomics have been rapidly developed. Currently, tandem mass spectrometry (MS/MS) is one of the key experimental tools for identification of structures of oligosaccharides. MS/MS can observe MS/MS peaks of fragmented glycan ions including cross-ring ions resulting from internal cleavages, which provide valuable information to infer glycan structures. Thus, the aim of de novo sequencing of glycans is to find the most probable assignments of observed MS/MS peaks to glycan substructures without databases. However, there are few satisfiable algorithms for glycan de novo sequencing from MS/MS spectra. We present a machine learning based approach to de novo sequencing of glycans from MS/MS spectrum. First, we build a suitable model for the fragmentation of glycans including cross-ring ions, and implement a solver that employs Lagrangian relaxation with a dynamic programming technique. Then, to optimize scores for the algorithm, we introduce a machine learning technique called structured support vector machines that enable us to learn parameters including scores for cross-ring ions from training data, i.e., known glycan mass spectra. Furthermore, we implement additional constraints for core structures of well-known glycan types including N-linked glycans and O-linked glycans. This enables us to predict more accurate glycan structures if the glycan type of given spectra is known. Computational experiments show that our algorithm performs accurate de novo sequencing of glycans. The implementation of our algorithm and the datasets are available at http://glyfon.dna.bio.keio.ac.jp/.