Computational Biology in the Cloud: Methods and New Insights from Computing at Scale

Computational Biology in the Cloud: Methods and New Insights from Computing at Scale
复制标题

云中的计算生物学:大规模计算的方法和新见解

DOI:
--
复制
发表时间:
2012
期刊:
Pacific Symposium on Biocomputing
影响因子:
--
通讯作者:
P. Kasson
P. Kasson
中科院分区:
--
文献类型:
--
作者:
P. Kasson

文献摘要

被引文献

相似文献

在过去的几年里,生物数据集的规模出现了爆炸性增长,新的、高度灵活的按需计算能力也在激增。基因组和元基因组测序、高通量蛋白质组学、关于分子结构和动力学的实验和模拟数据集提供的大量信息为极大地扩展洞察力提供了机会,但它也为千万亿级数据的计算、存储和解释带来了新的规模挑战。云计算资源具有帮助解决这些问题的潜力,因为它提供了一种计算和存储的实用模式:近乎无限的容量、突发使用的能力以及廉价而灵活的支付模式。云计算在大型生物数据集上的有效使用需要处理规模和健壮性等非同寻常的问题,因为当数据集增长10,000倍或更多时,性能限制因素可能会发生重大变化。因此,经常需要新的计算模式。云平台的使用还创造了新的机会来共享数据、减少重复,并通过使数据集和计算方法易于获得来提供易于重复性。
The past few years have seen both explosions in the size of biological data sets and the proliferation of new, highly flexible on-demand computing capabilities. The sheer amount of information available from genomic and metagenomic sequencing, high-throughput proteomics, experimental and simulation datasets on molecular structure and dynamics affords an opportunity for greatly expanded insight, but it creates new challenges of scale for computation, storage, and interpretation of petascale data. Cloud computing resources have the potential to help solve these problems by offering a utility model of computing and storage: near-unlimited capacity, the ability to burst usage, and cheap and flexible payment models. Effective use of cloud computing on large biological datasets requires dealing with non-trivial problems of scale and robustness, since performance-limiting factors can change substantially when a dataset grows by a factor of 10,000 or more. New computing paradigms are thus often needed. The use of cloud platforms also creates new opportunities to share data, reduce duplication, and to provide easy reproducibility by making the datasets and computational methods easily available.