Security Vulnerability Detection Using Deep Learning Natural Language Processing

Security Vulnerability Detection Using Deep Learning Natural Language Processing
复制标题

DOI:
10.1109/infocomwkshps51825.2021.9484500
复制
发表时间:
2021-05
期刊:
IEEE INFOCOM 2021 - IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS)
影响因子:
--
通讯作者:
Noah Ziems;Shaoen Wu
Noah Ziems;Shaoen Wu
中科院分区:
其他
文献类型:
--
作者:
Noah Ziems;Shaoen Wu

文献摘要

被引文献

相似文献

几十年来,在软件的安全漏洞被利用之前检测它们一直是一个具有挑战性的问题。传统的代码分析方法已经提出,但往往是无效的和低效的。在这项工作中,我们将软件漏洞检测建模为将源代码视为文本的自然语言处理(NLP)问题,并使用最近先进的深度学习NLP模型,在书面英语迁移学习的辅助下,解决自动软件漏洞检测问题。为了培训和测试,我们对NIST NVD/SARD数据库进行了预处理,用C编程语言构建了一个包含超过10万个文件的数据集,其中包含123种类型的漏洞。经过大量的实验,该方法在检测安全漏洞方面的准确率达到93%以上。
Detecting security vulnerabilities in software before they are exploited has been a challenging problem for decades. Traditional code analysis methods have been proposed, but are often ineffective and inefficient. In this work, we model software vulnerability detection as a natural language processing (NLP) problem with source code treated as texts, and address the auto-mated software venerability detection with recent advanced deep learning NLP models assisted by transfer learning on written English. For training and testing, we have preprocessed the NIST NVD/SARD databases and built a dataset of over 100,000 files in C programming language with 123 types of vulnerabilities. The extensive experiments generate the best performance of over 93% accuracy in detecting security vulnerabilities.