Emma in Action: Declarative Dataflows for Scalable Data Analysis

Emma in Action: Declarative Dataflows for Scalable Data Analysis
复制标题

DOI:
10.1145/2882903.2899396
复制
发表时间:
2016-06
期刊:
Proceedings of the 2016 International Conference on Management of Data
影响因子:
--
通讯作者:
Alexander B. Alexandrov;Andreas Salzmann;Georgi Krastev;Asterios Katsifodimos;V. Markl
Alexander B. Alexandrov;Andreas Salzmann;Georgi Krastev;Asterios Katsifodimos;V. Markl
中科院分区:
其他
文献类型:
--
作者:
Alexander B. Alexandrov;Andreas Salzmann;Georgi Krastev;Asterios Katsifodimos;V. Markl

文献摘要

被引文献

相似文献

基于二阶功能的并行数据流API最初被视为SQL的灵活替代方案。但是,随着时间的流逝,由于基础发动机必须暴露的物理方面数量,以促进有效的执行。为了保留足够的抽象水平,并降低了数据科学家的进入障碍,Spark和Flink等项目目前在其平行收集抽象的基础上提供了特定于域的API。该演示强调了基于深度语言嵌入的替代设计的好处。我们展示了艾玛(Emma) - 一种嵌入Scala中的编程语言。 Emma通过Scala的Forcrehensions(类似于SQL)等天然构造来促进并行收集处理。此外,Emma还提倡Quasi引用整个数据分析算法,而不是其单个数据流表达式。这允许将引用的代码分解为(顺序)控制流和(并行)数据流片段,优化上下文中的数据流,并将它们透明地将其卸载到Spark或Flink之类的引擎中。拟议的设计承诺,由于避免阻抗不匹配,从而提高了程序员的生产率,从而减少了数据分析的滞后时间和成本。
Parallel dataflow APIs based on second-order functions were originally seen as a flexible alternative to SQL. Over time, however, their complexity increased due to the number of physical aspects that had to be exposed by the underlying engines in order to facilitate efficient execution. To retain a sufficient level of abstraction and lower the barrier of entry for data scientists, projects like Spark and Flink currently offer domain-specific APIs on top of their parallel collection abstractions. This demonstration highlights the benefits of an alternative design based on deep language embedding. We showcase Emma - a programming language embedded in Scala. Emma promotes parallel collection processing through native constructs like Scala's for-comprehensions - a declarative syntax akin to SQL. In addition, Emma also advocates quasi-quoting the entire data analysis algorithm rather than its individual dataflow expressions. This allows for decomposing the quoted code into (sequential) control flow and (parallel) dataflow fragments, optimizing the dataflows in context, and transparently offloading them to an engine like Spark or Flink. The proposed design promises increased programmer productivity due to avoiding an impedance mismatch, thereby reducing the lag times and cost of data analysis.