Bigdata logs analysis based on seq2seq networks for cognitive Internet of Things

Bigdata logs analysis based on seq2seq networks for cognitive Internet of Things
复制标题

基于seq2seq网络的认知物联网大数据日志分析

DOI:
10.1016/j.future.2018.08.021
复制
发表时间:
2019
期刊:
Future Generation Computer Systems
影响因子:
--
通讯作者:
Patrick C.K. Hung
Patrick C.K. Hung
中科院分区:
其他
文献类型:
--
作者:
Pin Wu;Zhihui Lu;Quan Zhou;Zhidan Lei;Xiaoqiang Li;Meikang Qiu;Patrick C.K. Hung

文献摘要

参考文献

被引文献

相似文献

大数据系统在高速处理海量数据的同时,也会产生大量的日志。然而,人们很难根据海量、多源、异构的大数据日志来预测未来事件。本文提出了一种物联网(IoT)中海量日志智能计算和预测的综合方法。传统的机器学习、隐马尔可夫模型 (HMM) 和自回归积分移动平均模型 (ARIMA) 方法不够准确,无法预测基于时间序列的数据随时间的变化。在这项工作中,我们首先详细阐述了大数据日志的分布式收集和存储、事件定位以及矢量化表示。接下来,我们提出一种日志融合算法,通过去除噪声、添加时间戳和分类标签,将大数据每个组成部分的日志(非结构化文本数据)转换为结构化数据。然后,我们介绍了大数据系统的预测模型。我们使用注意力机制来改进序列到序列(seq2seq)算法,并添加一个调整器来全局拟合数据分布。我们的实验结果表明,用我们的方法训练的神经网络模型对于真实世界的数据具有良好的性能。与之前的预测方法相比,均方根误差(RMSE)降低了46.65%,R平方(R2)拟合度提高了14.28%。
While bigdata system processes high-volume data at high speed, it also generates a large amount of logs. However, it is hard for people to predict future events based on massive, multi-source, heterogeneous bigdata logs. This paper proposes a comprehensive method for smart computation and prediction of massive logs in the internet of things (IoT). Traditional machine learning, Hidden Markov Model (HMM) and Autoregressive Integrated Moving Average Model (ARIMA) methods are not accurate enough to predict time series based data over time. In this work we first elaborate the distributed collection and storage, event location, and vectorized representations of bigdata logs. Next, we present a log fusion algorithm to convert the logs (unstructured text data) of each component of bigdata into structured data by removing noise, adding timestamps and classification labels. Then, we introduce a predictive model for bigdata system. We use an attention mechanism to improve sequence to sequence (seq2seq) algorithm and add an adjustor to globally fit the data distribution. Our experimental results show that the neural network model trained by our method has a good performance with the real-world data. Compared with the previous predictive method, the root mean square error (RMSE) is reduced by 46.65% and the R-squared (R2) fitting degree is improved by 14.28%.
DOI: 10.3923/itj.2011.798.806
发表时间: 2011-04
期刊: Information Technology Journal
影响因子: --
作者:
Asif Iqbal Hajamydeen;N. Udzir;R. Mahmod;A. Ghani
通讯作者: Asif Iqbal Hajamydeen;N. Udzir;R. Mahmod;A. Ghani
DOI: 10.1145/2723576.2723581
发表时间: 2015-03
期刊: Proceedings of the Fifth International Conference on Learning Analytics And Knowledge
影响因子: --
作者:
Christopher A. Brooks;Craig D. S. Thompson;Stephanie D. Teasley
通讯作者: Christopher A. Brooks;Craig D. S. Thompson;Stephanie D. Teasley
DOI: --
发表时间: 2008-09
期刊: --
影响因子: --
作者:
R. Yusof;S. R. Selamat;S. Sahib
通讯作者: R. Yusof;S. R. Selamat;S. Sahib
实现自动日志解析以进行大规模日志数据分析
DOI: 10.1109/tdsc.2017.2762673
发表时间: 2018-11-01
影响因子: 7.3
作者:
He, Pinjia;Zhu, Jieming;Lyu, Michael R.
通讯作者: Lyu, Michael R.
DOI: 10.1109/mis.2016.34
发表时间: 2016-03
影响因子: 6.4
作者:
A. Sheth
通讯作者: A. Sheth