Advanced Analytics with the SAP HANA Database

Advanced Analytics with the SAP HANA Database
复制标题

使用 SAP HANA 数据库进行高级分析

DOI:
--
复制
发表时间:
2013
期刊:
International Conference on Data Technologies and Applications
影响因子:
--
通讯作者:
Norman May
Norman May
中科院分区:
--
文献类型:
--
作者:
Philippe Grosse;Wolfgang Lehner;Norman May

文献摘要

被引文献

相似文献

MapReduce作为一种编程范式提供了一个简单易用但非常强大的抽象,封装在两个二阶函数中:Map和Reduce。因此,它们允许定义单个顺序处理的任务,同时隐藏有关这些任务如何并行化和扩展的许多框架细节。在本文中,我们讨论了四个处理模式的背景下,分布式SAP HANA数据库,超越了经典的MapReduce范式。我们使用一些典型的机器学习算法来说明它们,并给出实验结果,演示数据流如何随并行任务的数量扩展。
MapReduce as a programming paradigm provides a simple-to-use yet very powerful abstraction encapsulated in two second-order functions: Map and Reduce. As such, they allow defining single sequentially processed tasks while at the same time hiding many of the framework details about how those tasks are parallelized and scaled out. In this paper we discuss four processing patterns in the context of the distributed SAP HANA database that go beyond the classic MapReduce paradigm. We illustrate them using some typical Machine Learning algorithms and present experimental results that demonstrate how the data flows scale out with the number of parallel tasks.