iProX in 2021: connecting proteomics data sharing with big data.

iProX in 2021: connecting proteomics data sharing with big data.
复制标题

2021年的iProX:连接蛋白质组学数据共享与大数据。

DOI:
10.1093/nar/gkab1081
复制
发表时间:
2022-01-07
影响因子:
14.9
通讯作者:
Zhu Y
Zhu Y
中科院分区:
生物学2区
文献类型:
--
作者:
Chen T;Ma J;Liu Y;Chen Z;Xiao N;Lu Y;Fu Y;Yang C;Li M;Wu S;Wang X;Li D;He F;Hermjakob H;Zhu Y

文献摘要

参考文献

被引文献

相似文献

蛋白质组学研究的快速发展产生了大量的实验数据。大数据平台的出现为处理这些大量数据提供了机会。集成蛋白质组资源iProX(https://www.iprox.cn)于2017年启动,通过2021年实施的最新大数据平台得到了极大的改进。这里,我们描述了自2019年在Nucleic Acids Research首次发表以来iProX的主要进展。首先,具有高可扩展性的超融合架构支持提交过程。一个hadoop集群可以存储大量的蛋白质组数据集,一个分布式的、RESTful风格的弹性搜索引擎可以在一秒钟内查询数百万条记录。此外,iProX 还添加了一些新功能,包括 ProteomeXchange 提出的通用频谱标识符 (USI) 机制、RESTful Web 服务 API 和高效再分析管道,以实现更好的开放数据共享。截至2021年8月,已向iProX提交了1526个数据集,总数据量达到92.42TB。随着大数据平台的实施,iProX可以支持PB级数据存储、千亿条谱记录以及秒级延迟的服务能力,满足快速发展的蛋白质组学领域的需求。
The rapid development of proteomics studies has resulted in large volumes of experimental data. The emergence of big data platform provides the opportunity to handle these large amounts of data. The integrated proteome resource, iProX (https://www.iprox.cn), which was initiated in 2017, has been greatly improved with an up-to-date big data platform implemented in 2021. Here, we describe the main iProX developments since its first publication in Nucleic Acids Research in 2019. First, a hyper-converged architecture with high scalability supports the submission process. A hadoop cluster can store large amounts of proteomics datasets, and a distributed, RESTful-styled Elastic Search engine can query millions of records within one second. Also, several new features, including the Universal Spectrum Identifier (USI) mechanism proposed by ProteomeXchange, RESTful Web Service API, and a high-efficiency reanalysis pipeline, have been added to iProX for better open data sharing. By the end of August 2021, 1526 datasets had been submitted to iProX, reaching a total data volume of 92.42TB. With the implementation of the big data platform, iProX can support PB-level data storage, hundreds of billions of spectra records, and second-level latency service capabilities that meet the requirements of the fast growing field of proteomics.
DOI: 10.1093/nar/gky869
发表时间: 2019-01-08
影响因子: 14.9
作者:
Ma J;Chen T;Wu S;Yang C;Bai M;Shu K;Li K;Zhang G;Jin Z;He F;Hermjakob H;Zhu Y
通讯作者: Zhu Y
DOI: 10.1093/nar/gky899
发表时间: 2019-01-08
影响因子: 14.9
作者:
Moriya Y;Kawano S;Okuda S;Watanabe Y;Matsumoto M;Takami T;Kobayashi D;Yamanouchi Y;Araki N;Yoshizawa AC;Tabata T;Iwasaki M;Sugiyama N;Tanaka S;Goto S;Ishihama Y
通讯作者: Ishihama Y
DOI: 10.1002/pmic.201100515
发表时间: 2012-04-01
期刊: PROTEOMICS
影响因子: 3.4
作者:
Farrah, Terry;Deutsch, Eric W.;Moritz, Robert L.
通讯作者: Moritz, Robert L.
在 HBase 中实现基于 XML 的海量生物数据管理
DOI: 10.1109/tcbb.2019.2915811
发表时间: 2020-11-01
影响因子: 4.5
作者:
Liu,Jian;Liu,Qiuru;Liu,Yongzhuang
通讯作者: Liu,Yongzhuang
DOI: 10.1038/s41592-021-01184-6
发表时间: 2021-07
期刊: Nature methods
影响因子: 48
作者:
Deutsch EW;Perez-Riverol Y;Carver J;Kawano S;Mendoza L;Van Den Bossche T;Gabriels R;Binz PA;Pullman B;Sun Z;Shofstahl J;Bittremieux W;Mak TD;Klein J;Zhu Y;Lam H;Vizcaíno JA;Bandeira N
通讯作者: Bandeira N