Big data analytics by automated generation of fuzzy rules for Network Forensics Readiness

Big data analytics by automated generation of fuzzy rules for Network Forensics Readiness
复制标题

DOI:
10.1016/j.asoc.2016.10.029
复制
发表时间:
2017-03
期刊:
Appl. Soft Comput.
影响因子:
--
通讯作者:
Andrii Shalaginov;K. Franke
Andrii Shalaginov;K. Franke
中科院分区:
其他
文献类型:
--
作者:
Andrii Shalaginov;K. Franke

文献摘要

被引文献

相似文献

分析网络取证中的大规模流量转储可能是一个复杂而重要的问题。这是收集证据和进行威胁情报以预见新的非法活动的重要一步。机器学习可以帮助自动支持取证专家的决策。此外,在实时系统中的应用可能会带来与取证准备和知识发现相关的额外障碍。我们相信,可以通过神经模糊的方法来缓解这种问题,神经模糊是人类可理解的模型和自动数据分析的融合。该方法包括用自组织特征映射对样本进行无监督的最优分组,并用人工神经网络调整模糊规则。在这项工作中,我们提出了改进的方法,使得以更快的方式提取更少的模糊规则成为可能。与现有的方法相比,新方法有两个优点。首先,我们改进了模糊斑块的估计方法。第二,通过合并额外的椭圆紧凑度信息来表示数据的参数化。利用椭圆旋转和奉承信息,可以得到隶属度函数。为了进一步增强该方法的泛化能力,在分组阶段对自举聚集进行了测试。最后,在具有500万个样本的入侵检测数据集上对该方法进行了测试,仅用了12条规则,分类正确率为94%。
Analysis of large-scale traffic dumps in Network Forensics can be a complex and non-trivial problem. This is an important step in collecting evidences and making threat intelligence to foresee new illegal activities. Machine Learning comes into help to automatically support decision of forensics expert. Furthermore, application in live systems may bring additional obstacles related to forensics readiness and knowledge discovery. We believe that it can be mitigated by means of Neuro-Fuzzy, a fusion of human-understandable model and automated data analytic. This method includes optimal unsupervised grouping of samples with so-called Self-Organizing Features Map and fuzzy rules tuning by Artificial Neural Network. In this work we propose improvements of the methods that makes it possible to extract fewer fuzzy rules in a faster manner. The new method has two advantages in comparison to existing. First, we improve the estimation of fuzzy patches. Second, parameterization that represents the data by incorporating additional ellipse compactness information. By using ellipse rotation and flattering information, the membership functions can be derived. To even further enhance the generalization of the method, the bootstrap aggregation was tested during the grouping phase. Finally, the method has been assessed on the intrusion detection dataset with a five millions samples with classification accuracy 94% using only 12 rules.