Parallel Framework for Data-Intensive Computing with XSEDE

Parallel Framework for Data-Intensive Computing with XSEDE
复制标题

DOI:
10.1145/3332186.3338097
复制
发表时间:
2019-07
期刊:
Proceedings of the Practice and Experience in Advanced Research Computing on Rise of the Machines (learning)
影响因子:
--
通讯作者:
R. Subramanian;Hui Zhang
R. Subramanian;Hui Zhang
中科院分区:
其他
文献类型:
--
作者:
R. Subramanian;Hui Zhang

文献摘要

相似文献

随着数据驱动分析的增加,对高性能计算资源的需求不断增加。有许多高性能计算中心为学术研究提供网络基础设施(CI)。然而,将这些资源提供给广泛的用户时存在访问障碍。刚接触数据分析领域的用户还没有能力利用 CI 提供的工具。在本文中,我们提出了一个框架,以降低向没有接受过使用 CI 功能培训的用户提供高性能计算资源时存在的访问障碍。该框架使用分而治之(DC)范式来执行数据密集型计算任务。它由三个主要组件组成 - 用户界面 (UI)、并行脚本生成器 (PSG) 和底层网络基础设施 (CI)。该框架的目标是提供一种用户友好的方法,以最少的用户干预并行化数据密集型计算任务。一些关键的设计目标是可用性、可扩展性和可重复性。用户可以专注于他们的问题,并将并行化细节留给框架。
With the increase in data-driven analytics, the demand for high performing computing resources has risen. There are many high-performance computing centers providing cyberinfrastructure (CI) for academic research. However, there exists access barriers in bringing these resources to a broad range of users. Users who are new to data analytics field are not yet equipped to take advantage of the tools offered by CI. In this paper, we propose a framework to lower the access barriers that exist in bringing the high-performance computing resources to users that do not have the training to utilize the capability of CI. The framework uses divide-and-conquer (DC) paradigm for data-intensive computing tasks. It consists of three major components - user interface (UI), parallel scripts generator (PSG) and underlying cyberinfrastructure (CI). The goal of the framework is to provide a user-friendly method for parallelizing data-intensive computing tasks with minimal user intervention. Some of the key design goals are usability, scalability and reproducibility. The users can focus on their problem and leave the parallelization details to the framework.