Synthesizing Cyber Intrusion Alerts using Generative Adversarial Networks

Synthesizing Cyber Intrusion Alerts using Generative Adversarial Networks
复制标题

使用生成对抗网络合成网络入侵警报

DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Christopher Sweet
Christopher Sweet
中科院分区:
--
文献类型:
--
作者:
Christopher Sweet

文献摘要

被引文献

相似文献

渗透企业计算机网络的网络攻击在数量、严重性和复杂性上持续增长,因为我们对此类网络的依赖性在增长。尽管如此,主动网络安全仍然是一个开放的挑战,因为网络警报数据往往无法用于研究。此外,可用的数据是随机分布的,不平衡的,缺乏同质性,并依赖于与网络结构的潜在方面的复杂交互。目前,还没有普遍接受的方法来建模和生成合成警报数据以供进一步研究;也没有指标来量化合成生成的警报的保真度或识别数据中的关键属性。这项工作提出了解决方案,网络警报的建模和如何评分的保真度,这样的模型。生成对抗网络被用来生成来自两个大学渗透测试比赛的网络警报数据。提供了定义网络警报数据度量的期望属性的标准列表。一些统计和信息理论的指标,如直方图交叉和条件熵,满足这些标准,并用于分析。使用这些度量,可以识别合成生成的警报的关键关系,并将其与来自地面实况分布的数据进行比较。最后,通过这些指标,我们表明,添加互信息约束模型的生成提高了输出的质量,并成功地捕捉警报发生的概率低。
Cyber attacks infiltrating enterprise computer networks continue to grow in number, severity, and complexity as our reliance on such networks grows. Despite this, proactive cyber security remains an open challenge as cyber alert data is often not available for study. Furthermore, the data that is available is stochastically distributed, imbalanced, lacks homogeneity, and relies on complex interactions with latent aspects of the network structure. Currently, there is no commonly accepted way to model and generate synthetic alert data for further study; there are also no metrics to quantify the fidelity of synthetically generated alerts or identify critical attributes within the data. This work proposes solutions to both the modeling of cyber alerts and how to score the fidelity of such models. Generative Adversarial Networks are employed to generate cyber alert data taken from two collegiate penetration testing competitions. A list of criteria defining desirable attributes for cyber alert data metrics is provided. Several statistical and information-theoretic metrics, such as histogram intersection and conditional entropy, meet these criteria and are used for analysis. Using these metrics, critical relationships of synthetically generated alerts may be identified and compared to data from the ground truth distribution. Finally, through these metrics, we show that adding a mutual information constraint to the model’s generation increases the quality of outputs and successfully captures alerts that occur with low probability.