An Automated Post-Mortem Analysis of Vulnerability Relationships using Natural Language Word Embeddings

An Automated Post-Mortem Analysis of Vulnerability Relationships using Natural Language Word Embeddings
复制标题

使用自然语言词嵌入对漏洞关系进行自动事后分析

DOI:
10.1016/j.procs.2021.04.018
复制
发表时间:
2021
期刊:
Procedia Computer Science
影响因子:
--
通讯作者:
Meneely, Andrew
Meneely, Andrew
中科院分区:
--
文献类型:
--
作者:
Meyers, Benjamin S.;Meneely, Andrew

文献摘要

相似文献

网络安全专家和软件工程师的日常活动-代码审查,问题跟踪,漏洞报告-不断为大量安全特定的自然语言做出贡献。在漏洞的情况下,了解其原因,后果和缓解措施对于从过去的错误中学习并在未来编写更好,更安全的代码至关重要。许多现有的漏洞评估方法,如CVSS,依赖于分类和数值指标来收集漏洞的见解,但这些工具无法捕捉漏洞之间的微妙复杂性和关系,因为它们没有检查开发人员留下的细微差别的自然语言工件。在这项工作中,我们希望发现意外的漏洞之间的关系,目的是改善目前的做法,事后分析的漏洞。为此,我们在两个语料库的漏洞描述从常见漏洞和暴露(CVE)和漏洞历史项目(VHP)的词嵌入模型进行了训练,执行层次凝聚聚类的词嵌入向量表示的整体语义含义的漏洞描述,并从漏洞集群的基础上,他们最常见的bigram的见解。我们发现:(1)具有相似后果和基于相似弱点的漏洞通常被聚集在一起,(2)聚类词嵌入识别出需要更详细描述的漏洞,(3)集群很少包含来自单个软件项目的漏洞。我们的方法是自动化的,可以很容易地应用到其他自然语言语料库。我们发布了我们工作中使用的所有语料库、模型和代码。
The daily activities of cybersecurity experts and software engineers—code reviews, issue tracking, vulnerability reporting—are constantly contributing to a massive wealth of security-specific natural language. In the case of vulnerabilities, understanding their causes, consequences, and mitigations is essential to learning from past mistakes and writing better, more secure code in the future. Many existing vulnerability assessment methodologies, like CVSS, rely on categorization and numerical metrics to glean insights into vulnerabilities, but these tools are unable to capture the subtle complexities and relationships between vulnerabilities because they do not examine the nuanced natural language artifacts left behind by developers. In this work, we want to discover unexpected relationships between vulnerabilities with the goal of improving upon current practices for post-mortem analysis of vulnerabilities. To that end, we trained word embedding models on two corpora of vulnerability descriptions from Common Vulnerabilities and Exposures (CVE) and the Vulnerability History Project (VHP), performed hierarchical agglomerative clustering on word embedding vectors representing the overall semantic meaning of vulnerability descriptions, and derived insights from vulnerability clusters based on their most common bigrams. We found that (1) vulnerabilities with similar consequences and based on similar weaknesses are often clustered together, (2) clustering word embeddings identified vulnerabilities that need more detailed descriptions, and (3) clusters rarely contained vulnerabilities from a single software project. Our methodology is automated and can be easily applied to other natural language corpora. We release all of the corpora, models, and code used in our work.