DarkEmbed: Exploit Prediction With Neural Language Models

DarkEmbed: Exploit Prediction With Neural Language Models
复制标题

DarkEmbed:利用神经语言模型进行预测

DOI:
10.1609/aaai.v32i1.11428
复制
发表时间:
2018
期刊:
The Journal of biological chemistry
影响因子:
--
通讯作者:
Kristina Lerman
Kristina Lerman
中科院分区:
--
文献类型:
--
作者:
N. Tavabi;Palash Goyal;Mohammed Almukaynizi;P. Shakarian;Kristina Lerman

文献摘要

被引文献

相似文献

软件漏洞可能会使计算机系统遭受恶意行为者的攻击。随着近年来发现的漏洞数量激增,为每个漏洞及时创建补丁并不总是可行的。同时,并非所有漏洞都会被攻击者利用;因此,通过评估漏洞被利用的可能性来确定漏洞的优先级已成为一个重要的研究问题。最近的工作使用机器学习技术通过分析社交媒体上有关漏洞的讨论来预测被利用的漏洞。这些方法依赖于传统的文本处理技术,该技术代表单词的统计特征,但无法捕获其上下文。为了应对这一挑战,我们提出了 DarkEmbed,这是一种神经语言建模方法,可以学习暗网/深网讨论的低维分布式表示(即嵌入),以预测漏洞是否会被利用。通过捕获人类语言的语言规律,例如句法、语义相似性和逻辑类比,学习的嵌入比传统的文本分析方法能够更好地对有关被利用漏洞的讨论进行分类。评估证明了学习嵌入对结构化文本(例如安全博客文章)和非结构化文本(暗网/深网帖子)的有效性。 DarkEmbed 在漏洞利用预测任务上的表现优于最先进的方法,F1 得分为 0.74。
Software vulnerabilities can expose computer systems to attacks by malicious actors. With the number of vulnerabilities discovered in the recent years surging, creating timely patches for every vulnerability is not always feasible. At the same time, not every vulnerability will be exploited by attackers; hence, prioritizing vulnerabilities by assessing the likelihood they will be exploited has become an important research problem. Recent works used machine learning techniques to predict exploited vulnerabilities by analyzing discussions about vulnerabilities on social media. These methods relied on traditional text processing techniques, which represent statistical features of words, but fail to capture their context. To address this challenge, we propose DarkEmbed, a neural language modeling approach that learns low dimensional distributed representations, i.e., embeddings, of darkweb/deepweb discussions to predict whether vulnerabilities will be exploited. By capturing linguistic regularities of human language, such as syntactic, semantic similarity and logic analogy, the learned embeddings are better able to classify discussions about exploited vulnerabilities than traditional text analysis methods. Evaluations demonstrate the efficacy of learned embeddings on both structured text (such as security blog posts) and unstructured text (darkweb/deepweb posts). DarkEmbed outperforms state-of-the-art approaches on the exploit prediction task with an F1-score of 0.74.