amulog: A General Log Analysis Framework for Diverse Template Generation Methods

amulog: A General Log Analysis Framework for Diverse Template Generation Methods
复制标题

DOI:
10.23919/cnsm50824.2020.9269049
复制
发表时间:
2020-11
期刊:
2020 16th International Conference on Network and Service Management (CNSM)
影响因子:
--
通讯作者:
Satoru Kobayashi;Y. Yamashiro;Kazuki Otomo;K. Fukuda
Satoru Kobayashi;Y. Yamashiro;Kazuki Otomo;K. Fukuda
中科院分区:
其他
文献类型:
--
作者:
Satoru Kobayashi;Y. Yamashiro;Kazuki Otomo;K. Fukuda

文献摘要

相似文献

分析来自大型IT系统的非结构化日志消息的方法之一是使用模板生成方法生成的日志模板对日志消息进行分类。但是,目前没有关于日志模板生成方法的比较和实际使用的共享知识,因为它们是在不同环境的基础上实施的。为此,我们设计并实现了一种适用于多种日志模板生成方法的通用日志分析框架Amulog。Amulog有三个关键功能:(1)将日志消息解析为头部和分段消息;(2)使用可伸缩的模板匹配方法对日志消息进行分类;(3)将结构化数据存储在数据库中。这个框架帮助我们轻松地利用与日志模板对应的时间序列数据进行进一步分析。我们使用从全国学术网络收集的日志数据集对amulog进行了评估,并证明了即使在超过10万个日志模板候选的情况下,它也可以在合理的时间内工作。
One of the ways to analyze unstructured log messages from large-scale IT systems is to classify log messages with log templates generated by template generation methods. However, there is currently no shared knowledge pertained to the comparison and practical use of log template generation methods because they are implemented on the basis of diverse environments. To this end, we design and implement amulog, a general log analysis framework for diverse log template generation methods. There are three key functions of amulog: (1) parsing log messages into headers and segmented messages, (2) classifying the log messages using a scalable template-matching method, and (3) storing the structured data in a database. This framework helps us easily utilize time-series data corresponding to the log templates for further analysis. We evaluate amulog with a log dataset collected from a nation-wide academic network and demonstrate that it works in a reasonable amount of time even with over 100,000 log template candidates.