Parallel Factor Analysis Enables Quantification and Identification of Highly Convolved Data-Independent-Acquired Protein Spectra.
Parallel Factor Analysis Enables Quantification and Identification of Highly Convolved Data-Independent-Acquired Protein Spectra.
复制标题
DOI:
10.1016/j.patter.2020.100137
复制
发表时间:
2020-12-11
期刊:
影响因子:
--
通讯作者:
Zelezniak A
中科院分区:
文献类型:
--
作者:
Buric F;Zrimec J;Zelezniak A
High-throughput data-independent acquisition (DIA) is the method of choice for quantitative proteomics, combining the best practices of targeted and shotgun approaches. The resultant DIA spectra are, however, highly convolved and with no direct precursor-fragment correspondence, complicating biological sample analysis. Here, we present CANDIA (canonical decomposition of data-independent-acquired spectra), a GPU-powered unsupervised multiway factor analysis framework that deconvolves multispectral scans to individual analyte spectra, chromatographic profiles, and sample abundances, using parallel factor analysis. The deconvolved spectra can be annotated with traditional database search engines or used as high-quality input for de novo sequencing methods. We demonstrate that spectral libraries generated with CANDIA substantially reduce the false discovery rate underlying the validation of spectral quantification. CANDIA covers up to 33 times more total ion current than library-based approaches, which typically use less than 5% of total recorded ions, thus allowing quantification and identification of signals from unexplored DIA spectra. Conventional DIA spectral libraries cover less than 3% of a scan's total ion count CANDIA deconvolves peptide signals by leveraging all scan data CANDIA uses GPUs to enable tensor algebra on massive DIA mass spectrometry data CANDIA output enables high-confidence and precise quantitative proteomics The latest high-throughput mass spectrometry-based technologies can record virtually all molecules from complex biological samples, providing a holistic picture of proteomes in cells and tissues and enabling an evaluation of the overall status of a person's health. However, current best practices are still only scratching the surface of the wealth of available information obtained from the massive proteome datasets, and efficient novel data-driven strategies are needed. Powered by advances in GPU hardware and open-source machine-learning frameworks, we developed a data-driven approach, CANDIA, which disassembles highly complex proteomics data into the elementary molecular signatures of the proteins in biological samples. Our work provides a performant and adaptable solution that complements existing mass spectrometry techniques. As the central mathematical methods are generic, other scientific fields that are dealing with highly convolved datasets will benefit from this work. We developed a software pipeline that enables deep analysis of very large proteomics data in real time, complementing existing techniques with unbiased unsupervised tensor decomposition.
登录
查看更多内容
影响因子:
4.4
作者:
Granholm, Viktor;Noble, William Stafford;Kall, Lukas
通讯作者:
Kall, Lukas
影响因子:
64.8
作者:
Harris CR;Millman KJ;van der Walt SJ;Gommers R;Virtanen P;Cournapeau D;Wieser E;Taylor J;Berg S;Smith NJ;Kern R;Picus M;Hoyer S;van Kerkwijk MH;Brett M;Haldane A;Del Río JF;Wiebe M;Peterson P;Gérard-Marchant P;Sheppard K;Reddy T;Weckesser W;Abbasi H;Gohlke C;Oliphant TE
通讯作者:
Oliphant TE
影响因子:
4.4
作者:
Deutsch EW;Perez-Riverol Y;Chalkley RJ;Wilhelm M;Tate S;Sachsenberg T;Walzer M;Käll L;Delanghe B;Böcker S;Schymanski EL;Wilmes P;Dorfer V;Kuster B;Volders PJ;Jehmlich N;Vissers JPC;Wolan DW;Wang AY;Mendoza L;Shofstahl J;Dowsey AW;Griss J;Salek RM;Neumann S;Binz PA;Lam H;Vizcaíno JA;Bandeira N;Röst H
通讯作者:
Röst H
影响因子:
7
作者:
Keller, Andrew;Bader, Samuel L.;Moritz, Robert L.
通讯作者:
Moritz, Robert L.
影响因子:
14.9
作者:
Gabriels, Ralf;Martens, Lennart;Degroeve, Sven
通讯作者:
Degroeve, Sven