Enabling User Driven Big Data Application on Remote Computing Resources

Enabling User Driven Big Data Application on Remote Computing Resources
复制标题

DOI:
10.1109/bigdata.2018.8622006
复制
发表时间:
2018-12
期刊:
2018 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Weijia Xu;Ruizhu Huang;Yige Wang
Weijia Xu;Ruizhu Huang;Yige Wang
中科院分区:
其他
文献类型:
--
作者:
Weijia Xu;Ruizhu Huang;Yige Wang

文献摘要

被引文献

相似文献

在计算资源需求的驱动下,数据驱动分析从本地计算资源向强大的远程计算资源(如云、高性能计算集群)迁移的需求越来越大。除了各种商业云服务之外,学术界也有丰富的高性能计算中心可供选择,提供网络基础设施(CI)产品。然而,在将这些资源提供给数据驱动的研究社区方面存在访问障碍。为了帮助降低这些访问障碍,提高远程资源对数据驱动分析的采用,我们提出了一种利用远程计算资源的新服务模型,使用户能够在远程计算资源上部署和运行其大数据应用程序作为web应用程序。该模型有几个关键的设计目标,包括支持交互性、可重用性和可再现性。与CI资源提供者通常支持的传统批处理模型相比,支持web应用程序接口可以实现交互式分析功能。用户利用一组预定义的任务模板通过配置文件设计应用程序,这些模板也可由用户扩展。从配置文件生成的应用程序是自包含的,可以在没有减轻系统特权的情况下进行部署。因此,可以用一种可以共享和重用的格式来描述和保存特别的分析例程。还可以通过配置文件描述和实现远程资源,以自动桥接应用程序与远程资源,并方便将来与不同资源的迁移。因此,可以通过配置文件保存分析任务,以实现再现性。在这里,我们详细介绍了我们提出的应用程序框架及其初步实现。我们通过聚合和分析实时tweet的实际用例演示了该框架的用法。
Driven by the computing resource requirement, there are increasing demands of migrating data driven analysis from local computing resource to powerful remote resources such as cloud and high performance computing cluster. In addition to various commercial cloud services, there are also rich selections of high performance computing centers in academia providing cyberinfrastructure (CI) offerings. However, access barriers exist in bring those resources to data driven research community at large. To help lower those access barriers and increase the adoption of utilization of remote resources for data driven analysis, we propose a new service model for utilizing remote computing resources, which empower users to deploy and run their big data application as a web application on remote computing resources. There are several key design goals of this model including enabling interactivity, reusability and reproducibility. Compare to the traditional batch-processing model commonly supported by CI resource providers, supporting a web application interface enables interactive analysis capabilities. Users design the application through a configuration file utilizing a set of predefined task templates that are also extensible by users. The application generated from the configuration file is self-contained and can be deployed without alleviated system privilege. Therefore, ad-hoc analysis routines can be described and preserved in a format that can be shared and re-used. Remote resources can also be described and implemented through configuration files to automatically bridge the application with remote resources and facilitate migration with different resources in the future. Consequently, analysis tasks can be preserved through the configuration file for reproducibility. Here we detail our proposed application framework and its preliminary implementations. We demonstrated usage of this framework with a practical use case of aggregating and analyzing live tweets.