Parallel Factor Analysis Enables Quantification and Identification of Highly Convolved Data-Independent-Acquired Protein Spectra.

Parallel Factor Analysis Enables Quantification and Identification of Highly Convolved Data-Independent-Acquired Protein Spectra.
复制标题

DOI:
10.1016/j.patter.2020.100137
复制
发表时间:
2020-12-11
期刊:
Patterns (New York, N.Y.)
影响因子:
--
通讯作者:
Zelezniak A
Zelezniak A
中科院分区:
其他
文献类型:
--
作者:
Buric F;Zrimec J;Zelezniak A

文献摘要

参考文献

被引文献

相似文献

高通量数据独立采集(DIA)是定量蛋白质组学的首选方法,结合了靶向和鸟枪法的最佳实践。然而,所得的DIA光谱是高度卷积的,并且没有直接的标记物-片段对应关系,使生物样品分析复杂化。在这里,我们提出了CANDIA(数据独立采集光谱的规范分解),一个GPU驱动的无监督多路因子分析框架,使用并行因子分析将多光谱扫描解卷积到单个分析物光谱,色谱图和样品丰度。去卷积的光谱可以用传统的数据库搜索引擎注释或用作从头测序方法的高质量输入。我们表明,与CANDIA产生的光谱库大大降低了错误的发现率的光谱量化的验证。CANDIA覆盖的总离子电流比基于库的方法多33倍,基于库的方法通常使用不到5%的总记录离子,从而允许对来自未探索的DIA光谱的信号进行定量和识别。传统的DIA光谱库覆盖不到扫描总离子计数的3% CANDIA通过利用所有扫描数据对肽信号进行反卷积CANDIA使用GPU对大量DIA质谱数据启用张量代数CANDIA输出实现高置信度和精确的定量蛋白质组学最新的高通量质谱技术可以记录复杂生物样品中的几乎所有分子,提供细胞和组织中蛋白质组的整体图像,并能够评估一个人的整体健康状况。然而,目前的最佳实践仍然只是从大量蛋白质组数据集获得的丰富可用信息的表面,需要有效的新数据驱动策略。在GPU硬件和开源机器学习框架的推动下,我们开发了一种数据驱动的方法CANDIA,该方法将高度复杂的蛋白质组学数据分解为生物样品中蛋白质的基本分子特征。我们的工作提供了一个性能和适应性强的解决方案,补充了现有的质谱技术。由于中心数学方法是通用的,其他处理高度卷积数据集的科学领域将从这项工作中受益。我们开发了一个软件管道,可以在真实的时间内对非常大的蛋白质组学数据进行深入分析,并通过无偏无监督张量分解补充现有技术。
High-throughput data-independent acquisition (DIA) is the method of choice for quantitative proteomics, combining the best practices of targeted and shotgun approaches. The resultant DIA spectra are, however, highly convolved and with no direct precursor-fragment correspondence, complicating biological sample analysis. Here, we present CANDIA (canonical decomposition of data-independent-acquired spectra), a GPU-powered unsupervised multiway factor analysis framework that deconvolves multispectral scans to individual analyte spectra, chromatographic profiles, and sample abundances, using parallel factor analysis. The deconvolved spectra can be annotated with traditional database search engines or used as high-quality input for de novo sequencing methods. We demonstrate that spectral libraries generated with CANDIA substantially reduce the false discovery rate underlying the validation of spectral quantification. CANDIA covers up to 33 times more total ion current than library-based approaches, which typically use less than 5% of total recorded ions, thus allowing quantification and identification of signals from unexplored DIA spectra. Conventional DIA spectral libraries cover less than 3% of a scan's total ion count CANDIA deconvolves peptide signals by leveraging all scan data CANDIA uses GPUs to enable tensor algebra on massive DIA mass spectrometry data CANDIA output enables high-confidence and precise quantitative proteomics The latest high-throughput mass spectrometry-based technologies can record virtually all molecules from complex biological samples, providing a holistic picture of proteomes in cells and tissues and enabling an evaluation of the overall status of a person's health. However, current best practices are still only scratching the surface of the wealth of available information obtained from the massive proteome datasets, and efficient novel data-driven strategies are needed. Powered by advances in GPU hardware and open-source machine-learning frameworks, we developed a data-driven approach, CANDIA, which disassembles highly complex proteomics data into the elementary molecular signatures of the proteins in biological samples. Our work provides a performant and adaptable solution that complements existing mass spectrometry techniques. As the central mathematical methods are generic, other scientific fields that are dealing with highly convolved datasets will benefit from this work. We developed a software pipeline that enables deep analysis of very large proteomics data in real time, complementing existing techniques with unbiased unsupervised tensor decomposition.
DOI: 10.1021/pr1012619
发表时间: 2011-05-06
影响因子: 4.4
作者:
Granholm, Viktor;Noble, William Stafford;Kall, Lukas
通讯作者: Kall, Lukas
DOI: 10.1038/s41586-020-2649-2
发表时间: 2020-09
期刊: Nature
影响因子: 64.8
作者:
Harris CR;Millman KJ;van der Walt SJ;Gommers R;Virtanen P;Cournapeau D;Wieser E;Taylor J;Berg S;Smith NJ;Kern R;Picus M;Hoyer S;van Kerkwijk MH;Brett M;Haldane A;Del Río JF;Wiebe M;Peterson P;Gérard-Marchant P;Sheppard K;Reddy T;Weckesser W;Abbasi H;Gohlke C;Oliphant TE
通讯作者: Oliphant TE
DOI: 10.1021/acs.jproteome.8b00485
发表时间: 2018-12-07
影响因子: 4.4
作者:
Deutsch EW;Perez-Riverol Y;Chalkley RJ;Wilhelm M;Tate S;Sachsenberg T;Walzer M;Käll L;Delanghe B;Böcker S;Schymanski EL;Wilmes P;Dorfer V;Kuster B;Volders PJ;Jehmlich N;Vissers JPC;Wolan DW;Wang AY;Mendoza L;Shofstahl J;Dowsey AW;Griss J;Salek RM;Neumann S;Binz PA;Lam H;Vizcaíno JA;Bandeira N;Röst H
通讯作者: Röst H
DOI: 10.1074/mcp.o114.044917
发表时间: 2015-05-01
影响因子: 7
作者:
Keller, Andrew;Bader, Samuel L.;Moritz, Robert L.
通讯作者: Moritz, Robert L.
DOI: 10.1093/nar/gkz299
发表时间: 2019-07-02
影响因子: 14.9
作者:
Gabriels, Ralf;Martens, Lennart;Degroeve, Sven
通讯作者: Degroeve, Sven