Democratizing data-independent acquisition proteomics analysis on public cloud infrastructures via the Galaxy framework.

Democratizing data-independent acquisition proteomics analysis on public cloud infrastructures via the Galaxy framework.
复制标题

DOI:
10.1093/gigascience/giac005
复制
发表时间:
2022-02-15
期刊:
影响因子:
9.2
通讯作者:
Schilling O
Schilling O
中科院分区:
生物学2区
文献类型:
--
作者:
Fahrner M;Föll MC;Grüning BA;Bernt M;Röst H;Schilling O

文献摘要

参考文献

被引文献

相似文献

数据独立采集(DIA)已成为全球质谱蛋白质组学研究的重要方法,因为它提供了对生物系统分子多样性的深入了解。然而,国防情报局的数据分析仍然具有挑战性,因为其复杂性高,数据和样本量大,需要专门的软件和庞大的计算基础设施。大多数可用的开源DIA软件需要基本的编程技能,并且仅涵盖完整DIA数据分析的一小部分。因此,DIA数据分析通常需要使用多种软件工具及其兼容性,严重限制了可用性和再现性。为了克服这一障碍,我们在Galaxy框架中集成了一套开源DIA工具,用于可重现和版本控制的数据处理。DIA套件包括OpenSwath、PyProphet、diapysef和swath2stats。我们已经编译了用于DIA处理的功能性Galaxy管道,它为这些预安装和预配置的工具提供了基于Web的图形用户界面,以便在Galaxy框架的免费访问,强大的计算资源上使用。除了共享原始数据和结果外,这种方法还可以通过完整配置实现无缝共享工作流程。我们证明了一个一体化的DIA管道在银河的尖峰案例研究数据集的分析的可用性。此外,还提供了广泛的培训材料,以进一步增加蛋白质组学社区的访问。在基于网络和用户友好的Galaxy框架中集成开源DIA分析套件,并结合广泛的培训材料,使广泛的研究社区能够进行可重复和透明的DIA数据分析。
Data-independent acquisition (DIA) has become an important approach in global, mass spectrometric proteomic studies because it provides in-depth insights into the molecular variety of biological systems. However, DIA data analysis remains challenging owing to the high complexity and large data and sample size, which require specialized software and vast computing infrastructures. Most available open-source DIA software necessitates basic programming skills and covers only a fraction of a complete DIA data analysis. In consequence, DIA data analysis often requires usage of multiple software tools and compatibility thereof, severely limiting the usability and reproducibility. To overcome this hurdle, we have integrated a suite of open-source DIA tools in the Galaxy framework for reproducible and version-controlled data processing. The DIA suite includes OpenSwath, PyProphet, diapysef, and swath2stats. We have compiled functional Galaxy pipelines for DIA processing, which provide a web-based graphical user interface to these pre-installed and pre-configured tools for their use on freely accessible, powerful computational resources of the Galaxy framework. This approach also enables seamless sharing workflows with full configuration in addition to sharing raw data and results. We demonstrate the usability of an all-in-one DIA pipeline in Galaxy by the analysis of a spike-in case study dataset. Additionally, extensive training material is provided to further increase access for the proteomics community. The integration of an open-source DIA analysis suite in the web-based and user-friendly Galaxy framework in combination with extensive training material empowers a broad community of researches to perform reproducible and transparent DIA data analysis.
使用 iRT(一种标准化保留时间)可以更有针对性地测量肽。
DOI: 10.1002/pmic.201100463
发表时间: 2012-04
期刊: PROTEOMICS
影响因子: 3.4
作者:
Escher, Claudia;Reiter, Lukas;MacLean, Brendan;Ossola, Reto;Herzog, Franz;Chilton, John;MacCoss, Michael J.;Rinner, Oliver
通讯作者: Rinner, Oliver
DOI: 10.1007/978-1-60761-444-9_22
发表时间: 2010
期刊: Methods in molecular biology (Clifton, N.J.)
影响因子: --
作者:
Deutsch, Eric W
通讯作者: Deutsch, Eric W
DOI: 10.1371/journal.pone.0153160
发表时间: 2016
期刊: PloS one
影响因子: 3.7
作者:
Blattmann P;Heusel M;Aebersold R
通讯作者: Aebersold R
DOI: 10.1021/acs.jproteome.1c00123
发表时间: 2021-06-21
影响因子: 4.4
作者:
Bichmann, Leon;Gupta, Shubham;Rost, Hannes
通讯作者: Rost, Hannes
DOI: 10.1074/mcp.ra119.001472
发表时间: 2019-10-01
影响因子: 7
作者:
Brenes, Alejandro;Hukelmann, Ens;Lamond, Angus, I
通讯作者: Lamond, Angus, I