Actor-based incremental tree data processing for large-scale machine learning applications

Actor-based incremental tree data processing for large-scale machine learning applications
复制标题

DOI:
10.1145/3358499.3361220
复制
发表时间:
2019-10
期刊:
Proceedings of the 9th ACM SIGPLAN International Workshop on Programming Based on Actors, Agents, and Decentralized Control
影响因子:
--
通讯作者:
K. Sakurai;Taiki Shimizu
K. Sakurai;Taiki Shimizu
中科院分区:
其他
文献类型:
--
作者:
K. Sakurai;Taiki Shimizu

文献摘要

相似文献

为了科普当前大规模数据集快速处理的需求,人们研究了许多基于树模型的在线机器学习技术。提出了一种增量式树数据处理的设计模式,即在内存中逐步构造按需树模型。我们的方法采用演员模型,利用多核和分布式计算机,而无需大量重写算法代码。该模式基本上将树中的节点定义为参与者,参与者是异步进程的单元,每个数据实例作为消息在参与者节点之间流动。具体研究了两种机器学习算法:VFDT决策树自顶向下生长算法和BIRCH层次聚类自底向上生长算法。为了支持VFDT,我们提出了一个复制根节点的扩展机制,这样它就可以解决瓶颈作为输入的开始。为了支持BIRCH,我们将递归构造过程分成异步步骤,通过遍历兄弟节点之间的额外水平链接来纠正目标节点。我们在Akka Java上实现了机器学习任务,并确认了大规模数据集任务的合理性能。
A number of online machine learning techniques based on tree model have been studied in order to cope with today's requirements of quickly processing large scale data-sets. We present a design pattern for incremental tree data processing as gradually constructing on-demand tree-model on memory. Our approach adopts the actor model as making use of multi-cores and distributed computers without largely rewriting code for algorithms. The pattern basically defines a node in the tree as an actor which is the unit of asynchronous processes and each data instance flows between actor nodes as a message. We study concrete two machine learning algorithms, VFDT for decision tree's top-down growth and BIRCH for hierarchical clustering's bottom up growth. For supporting VFDT, we propose an extension mechanism of replicating root nodes so that it can address bottleneck as starting of inputs. For supporting BIRCH, we split processes of recursive construction into asynchronous steps with correcting target node by traversing extra horizontal links between sibling nodes. We carried out machine learning tasks with our implementation on top of Akka Java, and we confirmed reasonable performance for the tasks with large scale data-sets.