Automated Generation and Selection of Interpretable Features for Enterprise Security

Automated Generation and Selection of Interpretable Features for Enterprise Security
复制标题

DOI:
10.1109/bigdata.2018.8621986
复制
发表时间:
2018-12
期刊:
2018 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Jiayi Duan;Ziheng Zeng;Alina Oprea;Shobha Vasudevan
Jiayi Duan;Ziheng Zeng;Alina Oprea;Shobha Vasudevan
中科院分区:
其他
文献类型:
--
作者:
Jiayi Duan;Ziheng Zeng;Alina Oprea;Shobha Vasudevan

文献摘要

相似文献

提出了一种有效的机器学习方法来检测企业安全日志中的恶意行为。我们的方法涉及特征工程,即通过对原始数据的特征应用运算符来生成新特征。我们从原始特征生成DNF公式,从中提取布尔函数,并利用傅立叶分析来生成新的奇偶特征,并根据它们的最高傅立叶系数对它们进行排序。我们在真实的企业数据集上演示了工程功能增强了广泛的分类器和聚类算法的性能。与原始数据特征分类相比,在牺牲不超过0.47%的准确率的情况下,工程特征的恶意召回率提高了50.6%。在对工程功能执行集群时,我们还观察到恶意集群的隔离效果更好。一般来说,根据我们感兴趣的度量标准,少数工程功能实现了比原始数据功能更高的性能。我们的功能工程方法还保留了可解释性,这是网络安全应用程序中的一个重要考虑因素。
We present an effective machine learning method for malicious activity detection in enterprise security logs. Our method involves feature engineering, or generating new features by applying operators on features of the raw data. We generate DNF formulas from raw features, extract Boolean functions from them, and leverage Fourier analysis to generate new parity features and rank them based on their highest Fourier coefficients. We demonstrate on real enterprise data sets that the engineered features enhance the performance of a wide range of classifiers and clustering algorithms. As compared to classification of raw data features, the engineered features achieve up to 50.6% improvement in malicious recall, while sacrificing no more than 0.47% in accuracy. We also observe better isolation of malicious clusters, when performing clustering on engineered features. In general, a small number of engineered features achieve higher performance than raw data features according to our metrics of interest. Our feature engineering method also retains interpretability, an important consideration in cyber security applications.