Use of an informed search space maximizes confidence of site-specific assignment of glycoprotein glycosylation.

Use of an informed search space maximizes confidence of site-specific assignment of glycoprotein glycosylation.
复制标题

使用知情的搜索空间可最大程度地提高糖蛋白糖基化特定于位点特异性分配的信心。

DOI:
10.1007/s00216-016-9970-5
复制
发表时间:
2017-01
影响因子:
4.3
通讯作者:
Zaia J
Zaia J
中科院分区:
化学2区
文献类型:
--
作者:
Khatri K;Klein JA;Zaia J

文献摘要

被引文献

相似文献

为了解释糖肽串联质谱学,有必要估计理论上的糖链组成和肽序列,称为搜索空间。要做到这一点,最简单的方法是根据公共数据库中的几组多糖成分建立一个天真的搜索空间,并假设目标糖蛋白是纯的。然而,纯化的糖蛋白通常含有共同纯化的糖蛋白污染物,这可能会混淆基于天真假设的串联质谱图的指定。此外,人们越来越需要从复杂的生物混合物中鉴定糖肽。幸运的是,用于糖组和蛋白质组学的液质联用(LC-MS)方法现在已经成熟并可用。我们证明了使用由测量的糖类和蛋白质组构建的知情搜索空间来定义用于解释糖蛋白质组学数据的搜索空间的价值。我们使用α-1-酸性糖蛋白混合到一组日益复杂的矩阵中来证明这一点。随着混合物复杂性的增加,天真的搜索空间气球和以可接受的置信度指定糖肽的能力减弱。此外,不可能识别没有被预测为天真搜索空间的一部分的糖肽。从已发布的葡聚糖糖组和蛋白质组学数据构建的搜索空间比其天真的对应数据要小,同时包括在混合物中检测到的所有蛋白质。这最大限度地提高了确定糖肽串联质谱图的能力。随着混合物复杂性的增加,每个糖肽前体离子的串联质谱数减少,导致目标糖蛋白的整体得分较低和覆盖深度减少。我们建议使用α-1-酸性糖蛋白作为衡量糖蛋白质组学研究的分析方法和生物信息学搜索参数的有效性的标准。
In order to interpret glycopeptide tandem mass spectra, it is necessary to estimate the theoretical glycan compositions and peptide sequences, known as the search space. The simplest way to do this is to build a naïve search space from sets of glycan compositions from public databases and to assume that the target glycoprotein is pure. Often, however, purified glycoproteins contain co-purified glycoprotein contaminants that have the potential to confound assignment of tandem mass spectra based on naïve assumptions. In addition, there is increasing need to characterize glycopeptides from complex biological mixtures. Fortunately, liquid chromatography-mass spectrometry (LC-MS) methods for glycomics and proteomics are now mature and accessible. We demonstrate the value of using an informed search space built from measured glycomes and proteomes to define the search space for interpretation of glycoproteomics data. We show this using α-1-acid glycoprotein (AGP) mixed into a set of increasingly complex matrices. As the mixture complexity increases, the naïve search space balloons and the ability to assign glycopeptides with acceptable confidence diminishes. In addition, it is not possible to identify glycopeptides not foreseen as part of the naïve search space. A search space built from released glycan glycomics and proteomics data is smaller than its naïve counterpart while including the full range of proteins detected in the mixture. This maximizes the ability to assign glycopeptide tandem mass spectra with confidence. As the mixture complexity increases, the number of tandem mass spectra per glycopeptide precursor ion decreases, resulting in lower overall scores and reduced depth of coverage for the target glycoprotein. We suggest use of α-1-acid glycoprotein as a standard to gauge effectiveness of analytical methods and bioinformatics search parameters for glycoproteomics studies.