DeepVD: Toward Class-Separation Features for Neural Network Vulnerability Detection

DeepVD: Toward Class-Separation Features for Neural Network Vulnerability Detection
复制标题

DOI:
10.1109/icse48619.2023.00189
复制
发表时间:
2023-05
期刊:
2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Wenbo Wang;Tien N. Nguyen;Shaohua Wang;Yi Li;Jiyuan Zhang;Aashish Yadavally
Wenbo Wang;Tien N. Nguyen;Shaohua Wang;Yi Li;Jiyuan Zhang;Aashish Yadavally
中科院分区:
其他
文献类型:
--
作者:
Wenbo Wang;Tien N. Nguyen;Shaohua Wang;Yi Li;Jiyuan Zhang;Aashish Yadavally

文献摘要

相似文献

包括深度学习(DL)在内的机器学习(ML)的进步已经使几种方法能够隐式地学习易受攻击的代码模式,以自动检测软件漏洞。最近的一项研究表明,尽管取得了成功,但现有的基于ML/DL的漏洞检测(VD)模型在区分两类漏洞和良性代码方面的能力有限。我们提出了DeepVD,一个基于图的神经网络VD模型,强调漏洞和良性代码之间的类分离特征。DeepVD在不同的抽象层次上利用了三种类型的类分离功能:语句类型(类似于词性标记),后支配树(涵盖常规的执行流)和异常流图(涵盖异常和错误处理流)。我们进行了几个实验,在包含303个项目的真实漏洞数据集中评估DeepVD,其中包含13,130个易受攻击的方法。我们的研究结果表明,DeepVD相对于最先进的基于ML/DL的VD,在精确度方面提高了13%-29.6%,在召回率方面提高了15.6%-28.9%,在F分数方面提高了16.4%-25.8%。我们的消融研究证实,我们设计的功能和组件有助于DeepVD实现漏洞和良性代码的高类分离性。
The advances of machine learning (ML) including deep learning (DL) have enabled several approaches to implicitly learn vulnerable code patterns to automatically detect software vulnerabilities. A recent study showed that despite successes, the existing ML/DL-based vulnerability detection (VD) models are limited in the ability to distinguish between the two classes of vulnerability and benign code. We propose DeepVD, a graph-based neural network VD model that emphasizes on class-separation features between vulnerability and benign code. DeepVDleverages three types of class-separation features at different levels of abstraction: statement types (similar to Part-of-Speech tagging), Post-Dominator Tree (covering regular flows of execution), and Exception Flow Graph (covering the exception and error-handling flows). We conducted several experiments to evaluate DeepVD in a real-world vulnerability dataset of 303 projects with 13,130 vulnerable methods. Our results show that DeepVD relatively improves over the state-of-the-art ML/DL-based VD approaches 13%–29.6% in precision, 15.6%–28.9% in recall, and 16.4%–25.8% in F-score. Our ablation study confirms that our designed features and components help DeepVDachieve high class-separability for vulnerability and benign code.