Deep Learning-based MSMS Spectra Reduction in Support of Running Multiple Protein Search Engines on Cloud.

Deep Learning-based MSMS Spectra Reduction in Support of Running Multiple Protein Search Engines on Cloud.
复制标题

DOI:
10.1109/bibm.2017.8217951
复制
发表时间:
2017-11
期刊:
Proceedings. IEEE International Conference on Bioinformatics and Biomedicine
影响因子:
--
通讯作者:
Gupta A
Gupta A
中科院分区:
其他
文献类型:
--
作者:
Maabreh M;Qolomany B;Alsmadi I;Gupta A

文献摘要

参考文献

相似文献

可用的蛋白质搜索引擎相对于所使用的匹配算法的多样性、其结果之间的低重叠率以及其覆盖范围的差异鼓励蛋白质组学界利用不同搜索引擎的整体解决方案。云计算技术的进步和分布式处理集群的可用性也可以为这一任务提供支持。然而,在这种情况下,数据传输和结果合并可能是主要瓶颈。数十亿个观测到的质谱、数百 GB 甚至可能是 TB 的数据的大量涌入,很容易导致拥塞,增加故障风险、性能不佳、增加更多计算成本并浪费可用资源。因此,在本研究中,我们提出了一种深度学习模型,以减轻云网络上的流量,从而降低云计算的成本。该模型取决于每个光谱的前 50 个强度及其 m/z 值,删除了预计不会通过参与搜索引擎多数投票的任何光谱。我们使用三个搜索引擎(即:pFind、Comet 和 X!Tandem)以及四个不同数据集的结果是有前景的,并促进了对深度学习的投资,以解决此类大数据问题。
The diversity of the available protein search engines with respect to the utilized matching algorithms, the low overlap ratios among their results and the disparity of their coverage encourage the community of proteomics to utilize ensemble solutions of different search engines. The advancing in cloud computing technology and the availability of distributed processing clusters can also provide support to this task. However, data transferring and results’ combining, in this case, could be the major bottleneck. The flood of billions of observed mass spectra, hundreds of Gigabytes or potentially Terabytes of data, could easily cause the congestions, increase the risk of failure, poor performance, add more computations’ cost, and waste available resources. Therefore, in this study, we propose a deep learning model in order to mitigate the traffic over cloud network and, thus reduce the cost of cloud computing. The model, which depends on the top 50 intensities and their m/z values of each spectrum, removes any spectrum which is predicted not to pass the majority voting of the participated search engines. Our results using three search engines namely: pFind, Comet and X!Tandem, and four different datasets are promising and promote the investment in deep learning to solve such type of Big data problems.
DOI: 10.1021/pr300694b
发表时间: 2012-12-01
影响因子: 4.4
作者:
Trudgian, David C.;Mirzaei, Hamid
通讯作者: Mirzaei, Hamid
DOI: 10.1021/ac0258709
发表时间: 2003-02-15
影响因子: 7.4
作者:
Fenyö, D;Beavis, RC
通讯作者: Beavis, RC
DOI: 10.1186/1471-2105-13-324
发表时间: 2012-12-05
期刊: BMC bioinformatics
影响因子: 3
作者:
Lewis S;Csordas A;Killcoyne S;Hermjakob H;Hoopmann MR;Moritz RL;Deutsch EW;Boyle J
通讯作者: Boyle J
DOI: 10.1074/mcp.o114.043380
发表时间: 2015-02-01
影响因子: 7
作者:
Slagel, Joseph;Mendoza, Luis;Moritz, Robert L.
通讯作者: Moritz, Robert L.
DOI: 10.1093/bioinformatics/bth947
发表时间: 2004-08-04
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bern, Marshall;Goldberg, David;Yates, John R., III
通讯作者: Yates, John R., III