High Performance Data Engineering Everywhere

High Performance Data Engineering Everywhere
复制标题

DOI:
10.1109/smds49396.2020.00022
复制
发表时间:
2020-07
期刊:
2020 IEEE International Conference on Smart Data Services (SMDS)
影响因子:
--
通讯作者:
Chathura Widanage;Niranda Perera;V. Abeykoon;Supun Kamburugamuve;Thejaka Amila Kanewala;Hasara Maithree;P. Wickramasinghe;A. Uyar;Gurhan Gunduz;G. Fox
Chathura Widanage;Niranda Perera;V. Abeykoon;Supun Kamburugamuve;Thejaka Amila Kanewala;Hasara Maithree;P. Wickramasinghe;A. Uyar;Gurhan Gunduz;G. Fox
中科院分区:
其他
文献类型:
--
作者:
Chathura Widanage;Niranda Perera;V. Abeykoon;Supun Kamburugamuve;Thejaka Amila Kanewala;Hasara Maithree;P. Wickramasinghe;A. Uyar;Gurhan Gunduz;G. Fox

文献摘要

相似文献

机器和深度学习领域取得的惊人进步是大数据时代企业和研究社区的一大亮点。现代应用程序需要的资源超出了单个节点的能力。然而,这只是整个数据处理环境所面临的问题的一小部分,它还必须支持大量的数据工程,以进行前后数据处理、通信和系统集成。数据分析工具的一个重要要求是能够轻松地与多种语言的现有框架集成,从而提高用户的生产力和效率。所有这些都需要一种高效且高度分布式的数据处理集成方法,但当今许多流行的数据分析工具无法同时满足所有这些要求。在本文中,我们介绍了Cylon,这是一个开源的高性能分布式数据处理库,可以与现有的大数据和AI/ML框架无缝集成。它是在紧凑的数据结构之上使用灵活的C++核心开发的,并公开了与C++,Java和Python的语言绑定。我们详细讨论了Cylon的架构,并揭示了如何将其作为库导入现有应用程序或作为独立框架运行。最初的实验表明,Cylon增强了流行的工具,如Apache Spark和Dask,在关键操作和更好的组件链接方面具有重大的性能改进。最后,我们展示了它的设计如何使Cylon能够以最小的开销跨平台使用,其中包括流行的AI工具,如PyTorch,Tensorflow和Quixyter notebook。
The amazing advances being made in the fields of machine and deep learning are a highlight of the Big Data era for both enterprise and research communities. Modern applications require resources beyond a single node's ability to provide. However this is just a small part of the issues facing the overall data processing environment, which must also support a raft of data engineering for pre- and post-data processing, communication, and system integration. An important requirement of data analytics tools is to be able to easily integrate with existing frameworks in a multitude of languages, thereby increasing user productivity and efficiency. All this demands an efficient and highly distributed integrated approach for data processing, yet many of today's popular data analytics tools are unable to satisfy all these requirements at the same time. In this paper we present Cylon, an open-source high performance distributed data processing library that can be seamlessly integrated with existing Big Data and AI/ML frameworks. It is developed with a flexible C++ core on top of a compact data structure and exposes language bindings to C++, Java, and Python. We discuss Cylon's architecture in detail, and reveal how it can be imported as a library to existing applications or operate as a standalone framework. Initial experiments show that Cylon enhances popular tools such as Apache Spark and Dask with major performance improvements for key operations and better component linkages. Finally, we show how its design enables Cylon to be used cross-platform with minimum overhead, which includes popular AI tools such as PyTorch, Tensorflow, and Jupyter notebooks.