Copernicus, a hybrid dataflow and peer-to-peer scientific computing platform for efficient large-scale ensemble sampling

Copernicus, a hybrid dataflow and peer-to-peer scientific computing platform for efficient large-scale ensemble sampling
复制标题

Copernicus,一种混合​​数据流和点对点科学计算平台,用于高效的大规模集成采样

DOI:
10.1016/j.future.2016.11.004
复制
发表时间:
2017
期刊:
Future Gener. Comput. Syst.
影响因子:
--
通讯作者:
E. Lindahl
E. Lindahl
中科院分区:
--
文献类型:
--
作者:
Iman Pouya;S. Pronk;M. Lundborg;E. Lindahl

文献摘要

参考文献

被引文献

相似文献

计算密集型应用已逐渐从大规模并行的超级计算机转向按需获取的容量资源。尤其是云计算和MapReduce在工业中的大规模采用,而传统的高性能计算(HPC)在科学和工程计算中的使用已经很难利用这种类型的资源。然而,随着并行性的增强而不是更快的处理器,越来越多的应用程序已经在算法级别上使用基于采样和集成的松散耦合的方法来实现并行性。虽然这些不能简单地表示为MapReduce,但它们高度依赖于吞吐量计算。有许多通用和强大的框架,但特别是对于科学计算中基于采样的算法,拥有高度了解潜在物理问题的平台和调度器具有一些明显的优势。在这里,我们介绍如何结合数据流编程、点对点技术和点对点网络来解决这些挑战。这允许以采样为重点的工作流、任务生成、依赖关系跟踪的自动化,尤其是将这些分发到从超级计算机到云和分布式计算(跨防火墙和脆弱网络)的各种计算资源集。工作流是使用现有程序从模块定义的,这使得它们可以在不需要编程的情况下重用。该系统通过透明地处理节点故障来实现弹性,同时将由于检查点而造成的计算时间损失降至最低,并且单个服务器可以管理数十万个核心,例如用于计算化学应用。
Compute-intensive applications have gradually changed focus from massively parallel supercomputers to capacity as a resource obtained on-demand. This is particularly true for the large-scale adoption of cloud computing and MapReduce in industry, while it has been difficult for traditional high-performance computing (HPC) usage in scientific and engineering computing to exploit this type of resources. However, with the strong trend of increasing parallelism rather than faster processors, a growing number of applications target parallelism already on the algorithm level with loosely coupled approaches based on sampling and ensembles. While these cannot trivially be formulated as MapReduce, they are highly amenable to throughput computing. There are many general and powerful frameworks, but in particular for sampling-based algorithms in scientific computing there are some clear advantages from having a platform and scheduler that are highly aware of the underlying physical problem. Here, we present how these challenges are addressed with combinations of dataflow programming, peer-to-peer techniques and peer-to-peer networks in theCopernicusplatform. This allows automation of sampling-focused workflows, task generation, dependency tracking, and not least distributing these to a diverse set of compute resources ranging from supercomputers to clouds and distributed computing (across firewalls and fragile networks). Workflows are defined from modules using existing programs, which makes them reusable without programming requirements. The system achieves resiliency by handling node failures transparently with minimal loss of computing time due to checkpointing, and a single server can manage hundreds of thousands of cores e.g. for computational chemistry applications.
DOI: 10.1021/jp0777059
发表时间: 2008-03-20
影响因子: 3.3
作者:
Pan, Albert C.;Sezer, Deniz;Roux, Benoit
通讯作者: Roux, Benoit
DOI: 10.1063/1.3565032
发表时间: 2011-05-07
影响因子: 4.4
作者:
Prinz, Jan-Hendrik;Wu, Hao;Noe, Frank
通讯作者: Noe, Frank
DOI: 10.1063/1.1738640
发表时间: 2004-06-15
影响因子: 4.4
作者:
Faradjian, AK;Elber, R
通讯作者: Elber, R