Improving Proteoform Identifications in Complex Systems Through Integration of Bottom-Up and Top-Down Data

Improving Proteoform Identifications in Complex Systems Through Integration of Bottom-Up and Top-Down Data
复制标题

DOI:
10.1021/acs.jproteome.0c00332
复制
发表时间:
2020-08-07
影响因子:
4.4
通讯作者:
Smith, Lloyd M.
Smith, Lloyd M.
中科院分区:
生物学2区
文献类型:
--
作者:
Schaffer, Leah, V;Millikin, Robert J.;Smith, Lloyd M.

文献摘要

被引文献

相似文献

细胞功能是由一组巨大而多样的蛋白质形式来执行的。蛋白质形式是由于遗传变异、RNA剪接和翻译后修饰(PTM)而产生的特定形式的蛋白质。对完整蛋白质的自上而下的质谱分析使蛋白质形态的鉴定成为可能,包括来自序列切割事件或含有多个PTM的蛋白质形态。相反,自下而上的蛋白质组学识别多肽,这需要蛋白质推断,而不会产生蛋白质形式的鉴定。我们在这里寻求利用这两种数据类型之间的协同效应来提高整体蛋白质组分析的质量和深度。为此,我们在软件程序Proteoform Suite中自动化了多酶自下而上和自上而下分析结果的大规模集成,并将其应用于人类Jurkat T淋巴细胞系的蛋白质组分析。我们在Proteoform Suite中实现了最近开发的用于自上而下串联质谱学(MS/MS)鉴定的蛋白质形式水平分类方案,该方案使用户能够观察每个蛋白质形式鉴定的模糊程度和类型,包括哪些模糊的蛋白质形式鉴定得到自下而上水平的证据支持。我们使用Proteoform Suite从自下而上的分析中找到了自上而下的识别有助于蛋白质推断的实例,反之,自下而上的多肽识别有助于蛋白质形式的PTM定位。我们还展示了使用自下而上的数据来推断样品中可能存在的蛋白质候选,允许通过对MS1光谱的完整质量分析来确认这种蛋白质候选。在免费提供的软件程序Proteoform Suite中实现这些功能使用户能够整合大规模的自上而下和自下而上的数据集,并利用它们之间的协同作用来改进和扩展蛋白质组分析。
Cellular functions are performed by a vast and diverse set of proteoforms. Proteoforms are the specific forms of proteins produced as a result of genetic variations, RNA splicing, and post-translational modifications (PTMs). Top-down mass spectrometric analysis of intact proteins enables proteoform identification, including proteoforms derived from sequence cleavage events or harboring multiple PTMs. In contrast, bottom-up proteomics identifies peptides, which necessitates protein inference and does not yield proteoform identifications. We seek here to exploit the synergies between these two data types to improve the quality and depth of the overall proteomic analysis. To this end, we automated the large-scale integration of results from multiprotease bottom-up and top-down analyses in the software program Proteoform Suite and applied it to the analysis of proteoforms from the human Jurkat T lymphocyte cell line. We implemented the recently developed proteoform-level classification scheme for top-down tandem mass spectrometry (MS/MS) identifications in Proteoform Suite, which enables users to observe the level and type of ambiguity for each proteoform identification, including which of the ambiguous proteoform identifications are supported by bottom-up-level evidence. We used Proteoform Suite to find instances where top-down identifications aid in protein inference from bottom-up analysis and conversely where bottom-up peptide identifications aid in proteoform PTM localization. We also show the use of bottom-up data to infer proteoform candidates potentially present in the sample, allowing confirmation of such proteoform candidates by intact-mass analysis of MS1 spectra. The implementation of these capabilities in the freely available software program Proteoform Suite enables users to integrate large-scale top-down and bottom-up data sets and to utilize the synergies between them to improve and extend the proteomic analysis.