Implementing the reuse of public DIA proteomics datasets: from the PRIDE database to Expression Atlas.

Implementing the reuse of public DIA proteomics datasets: from the PRIDE database to Expression Atlas.
复制标题

DOI:
10.1038/s41597-022-01380-9
复制
发表时间:
2022-06-14
期刊:
影响因子:
9.8
通讯作者:
Vizcaino, Juan Antonio
Vizcaino, Juan Antonio
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Walzer, Mathias;Garcia-Seisdedos, David;Prakash, Ananth;Brack, Paul;Crowther, Peter;Graham, Robert L.;George, Nancy;Mohammed, Suhaib;Moreno, Pablo;Papatheodorou, Irene;Hubbard, Simon J.;Vizcaino, Juan Antonio

文献摘要

参考文献

被引文献

相似文献

公共领域中基于质谱(MS)的蛋白质组学数据集的数量不断增加,特别是那些由数据独立采集(DIA)方法(如SWATH-MS)生成的数据。与数据依赖采集数据集不同,DIA数据集的重复使用迄今为止相当有限,尽管其潜力很大,但由于所涉及的技术挑战。我们为公共SWATH-MS数据集引入了一个(重新)分析管道,其中包括元数据注释协议,MS数据分析,统计分析的自动化工作流程以及将结果集成到Expression Atlas资源中的组合。自动化与Nextflow协调,使用容器化的开放分析软件工具,使管道随时可用和可重现。为了证明其实用性,我们重新分析了PRIDE数据库中的10个公共DIA数据集,包括1,278次SWATH-MS运行。评价了分析的稳健性,并将结果与原始出版物中获得的结果进行了比较。最终表达值被整合到Expression Atlas中,使SWATH-MS实验更广泛地可用,并将其与来自其他蛋白质组学和转录组学数据集的表达数据相结合。
The number of mass spectrometry (MS)-based proteomics datasets in the public domain keeps increasing, particularly those generated by Data Independent Acquisition (DIA) approaches such as SWATH-MS. Unlike Data Dependent Acquisition datasets, the re-use of DIA datasets has been rather limited to date, despite its high potential, due to the technical challenges involved. We introduce a (re-)analysis pipeline for public SWATH-MS datasets which includes a combination of metadata annotation protocols, automated workflows for MS data analysis, statistical analysis, and the integration of the results into the Expression Atlas resource. Automation is orchestrated with Nextflow, using containerised open analysis software tools, rendering the pipeline readily available and reproducible. To demonstrate its utility, we reanalysed 10 public DIA datasets from the PRIDE database, comprising 1,278 SWATH-MS runs. The robustness of the analysis was evaluated, and the results compared to those obtained in the original publications. The final expression values were integrated into Expression Atlas, making SWATH-MS experiments more widely available and combining them with expression data originating from other proteomics and transcriptomics datasets.
DOI: 10.1038/nm.3807
发表时间: 2015-04
期刊: Nature medicine
影响因子: 82.9
作者:
通讯作者: --
DOI: 10.1038/s41598-018-26170-5
发表时间: 2018-05-21
期刊: Scientific reports
影响因子: 4.6
作者:
Talavera D;Kershaw CJ;Costello JL;Castelli LM;Rowe W;Sims PFG;Ashe MP;Grant CM;Pavitt GD;Hubbard SJ
通讯作者: Hubbard SJ
DOI: 10.1016/j.jprot.2019.03.005
发表时间: 2019-05-30
影响因子: 3.3
作者:
He, Bing;Shi, Jim;Zhu, Hao-Jie
通讯作者: Zhu, Hao-Jie
使用 iRT(一种标准化保留时间)可以更有针对性地测量肽。
DOI: 10.1002/pmic.201100463
发表时间: 2012-04
期刊: PROTEOMICS
影响因子: 3.4
作者:
Escher, Claudia;Reiter, Lukas;MacLean, Brendan;Ossola, Reto;Herzog, Franz;Chilton, John;MacCoss, Michael J.;Rinner, Oliver
通讯作者: Rinner, Oliver
蛋白质组学样本元数据表示多组学集成和大数据分析。
DOI: 10.1038/s41467-021-26111-3
发表时间: 2021-10-06
影响因子: 16.6
作者:
Dai C;Füllgrabe A;Pfeuffer J;Solovyeva EM;Deng J;Moreno P;Kamatchinathan S;Kundu DJ;George N;Fexova S;Grüning B;Föll MC;Griss J;Vaudel M;Audain E;Locard-Paulet M;Turewicz M;Eisenacher M;Uszkoreit J;Van Den Bossche T;Schwämmle V;Webel H;Schulze S;Bouyssié D;Jayaram S;Duggineni VK;Samaras P;Wilhelm M;Choi M;Wang M;Kohlbacher O;Brazma A;Papatheodorou I;Bandeira N;Deutsch EW;Vizcaíno JA;Bai M;Sachsenberg T;Levitsky LI;Perez-Riverol Y
通讯作者: Perez-Riverol Y