Geelytics: Enabling On-Demand Edge Analytics over Scoped Data Sources

Geelytics: Enabling On-Demand Edge Analytics over Scoped Data Sources
复制标题

Geelytics:在范围数据源上实现按需边缘分析

DOI:
--
复制
发表时间:
2016
期刊:
BigData Congress [Services Society]
影响因子:
--
通讯作者:
M. Bauer
M. Bauer
中科院分区:
--
文献类型:
--
作者:
Bin Cheng;Apostolos Papageorgiou;M. Bauer

文献摘要

被引文献

相似文献

大规模物联网(IoT)系统通常由大量地理上分布在物理环境中的传感器和执行器组成。为了对真实的情况做出快速反应,通常需要通过接近物联网设备的实时流处理来桥接传感器和执行器。Apache Storm和S4等现有流处理平台专为集群或云中的密集流处理而设计,但它们不适合大规模物联网系统,其中处理任务预计由执行器按需触发,然后在云边缘环境中分配和执行。为了填补这一空白,我们设计并实现了一个名为Geelytics的新系统,该系统可以通过传感器和执行器的物联网友好接口,在范围数据源上实现按需边缘分析。本文介绍了它的设计,实现,接口和核心算法。已经构建了三个示例应用程序,以展示Geelytics在实现高级物联网边缘分析方面的潜力。我们的初步评估结果表明,我们可以减少99%的带宽成本在人脸检测的例子,实现不到10毫秒的反应延迟和约1.5秒的启动延迟在离群值检测的例子,也节省了65%的重复计算成本通过共享中间结果在数据聚合的例子。
Large-scale Internet of Things (IoT) systems typically consist of a large number of sensors and actuators distributed geographically in a physical environment. To react fast on real time situations, it is often required to bridge sensors and actuators via real-time stream processing close to IoT devices. Existing stream processing platforms like Apache Storm and S4 are designed for intensive stream processing in a cluster or in the Cloud, but they are unsuitable for large scale IoT systems in which processing tasks are expected to be triggered by actuators on-demand and then be allocated and performed in a Cloud-Edge environment. To fill this gap, we designed and implemented a new system called Geelytics, which can enable on-demand edge analytics over scoped data sources via IoT-friendly interfaces to sensors and actuators. This paper presents its design, implementation, interfaces, and core algorithms. Three example applications have been built to showcase the potential of Geelytics in enabling advanced IoT edge analytics. Our preliminary evaluation results demonstrate that we can reduce the bandwidth cost by 99% in a face detection example, achieve less than 10 milliseconds reacting latency and about 1.5 seconds startup latency in an outlier detection example, and also save 65% duplicated computation cost via sharing intermediate results in a data aggregation example.