Towards Automated Log Parsing for Large-Scale Log Data Analysis

Towards Automated Log Parsing for Large-Scale Log Data Analysis
复制标题

实现自动日志解析以进行大规模日志数据分析

DOI:
10.1109/tdsc.2017.2762673
复制
发表时间:
2018-11-01
影响因子:
7.3
通讯作者:
Lyu, Michael R.
Lyu, Michael R.
中科院分区:
计算机科学2区
文献类型:
--
作者:
He, Pinjia;Zhu, Jieming;Lyu, Michael R.

文献摘要

被引文献

相似文献

日志广泛用于系统管理中以保证可靠性,因为它们通常是记录生产中详细系统运行时行为的唯一可用数据。由于日志的大小不断增加,开发人员(和运营商)希望通过应用数据挖掘方法来自动化分析,因此需要结构化输入数据(例如矩阵)。这引发了许多关于日志解析的研究,旨在将自由文本日志消息转换为结构化事件。然而,由于缺乏这些日志解析器的开源实现和性能比较基准,开发人员在将其应用于实践时不太可能意识到现有日志解析器的有效性及其局限性。他们必须经常重新实现或重新设计一个,这是耗时且多余的。在本文中,我们首先对当前最先进的日志解析器进行了表征研究,并评估了它们在包含超过一千万条日志消息的五个真实数据集上的功效。我们确定,尽管这些解析器的整体准确性很高,但它们在所有数据集上并不稳健。当日志增长到大规模(例如,2 亿条日志消息)时(这在实践中很常见),这些解析器的效率不足以在单台计算机上处​​理此类数据。为了解决上述限制,我们在大规模数据处理平台 Spark 之上设计并实现了并行日志解析器(即 POP)。已经进行了综合实验来评估合成数据集和真实数据集上的 POP。评估结果证明了POP在后续日志挖掘任务中的准确性、效率和有效性方面的能力。
Logs are widely used in system management for dependability assurance because they are often the only data available that record detailed system runtime behaviors in production. Because the size of logs is constantly increasing, developers (and operators) intend to automate their analysis by applying data mining methods, therefore structured input data (e.g., matrices) are required. This triggers a number of studies on log parsing that aims to transform free-text log messages into structured events. However, due to the lack of open-source implementations of these log parsers and benchmarks for performance comparison, developers are unlikely to be aware of the effectiveness of existing log parsers and their limitations when applying them into practice. They must often reimplement or redesign one, which is time-consuming and redundant. In this paper, we first present a characterization study of the current state of the art log parsers and evaluate their efficacy on five real-world datasets with over ten million log messages. We determine that, although the overall accuracy of these parsers is high, they are not robust across all datasets. When logs grow to a large scale (e.g., 200 million log messages), which is common in practice, these parsers are not efficient enough to handle such data on a single computer. To address the above limitations, we design and implement a parallel log parser (namely POP) on top of Spark, a large-scale data processing platform. Comprehensive experiments have been conducted to evaluate POP on both synthetic and real-world datasets. The evaluation results demonstrate the capability of POP in terms of accuracy, efficiency, and effectiveness on subsequent log mining tasks.