Automated identification of security issues from commit messages and bug reports

Automated identification of security issues from commit messages and bug reports
复制标题

DOI:
10.1145/3106237.3117771
复制
发表时间:
2017-08
期刊:
Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering
影响因子:
--
通讯作者:
Yaqin Zhou;Asankhaya Sharma
Yaqin Zhou;Asankhaya Sharma
中科院分区:
其他
文献类型:
--
作者:
Yaqin Zhou;Asankhaya Sharma

文献摘要

被引文献

相似文献

开源库中的漏洞数量正在迅速增加。然而,其中大多数都没有公开披露。这些未识别的漏洞使开发人员的产品面临被黑客攻击的风险,因为他们越来越依赖开源库来快速组装和构建软件。为了在开源库中找到未识别的漏洞并确保现代软件开发的安全,我们描述了一个高效的自动漏洞识别系统,该系统采用自然语言处理和机器学习技术,面向真实的实时跟踪大型项目。基于使用GitHub、JIRA和Bugzilla的开源项目中提交消息和错误报告的潜在信息,我们的K-fold堆叠分类器在漏洞识别方面取得了可喜的成果。与现有的基于SVM的分类器在提交消息中的漏洞识别方面的工作相比,我们在保持相同召回率的情况下将精度提高了54.55%。对于错误报告,我们实现了更高的精度为0.70和召回率为0.71相比,现有的工作。此外,在SourceClear上运行训练模型超过3个月的观察结果显示,精确率为0.83,召回率为0.74,并检测到349个隐藏漏洞,证明了所提出方法的有效性和通用性。
The number of vulnerabilities in open source libraries is increasing rapidly. However, the majority of them do not go through public disclosure. These unidentified vulnerabilities put developers' products at risk of being hacked since they are increasingly relying on open source libraries to assemble and build software quickly. To find unidentified vulnerabilities in open source libraries and secure modern software development, we describe an efficient automatic vulnerability identification system geared towards tracking large-scale projects in real time using natural language processing and machine learning techniques. Built upon the latent information underlying commit messages and bug reports in open source projects using GitHub, JIRA, and Bugzilla, our K-fold stacking classifier achieves promising results on vulnerability identification. Compared to the state of the art SVM-based classifier in prior work on vulnerability identification in commit messages, we improve precision by 54.55% while maintaining the same recall rate. For bug reports, we achieve a much higher precision of 0.70 and recall rate of 0.71 compared to existing work. Moreover, observations from running the trained model at SourceClear in production for over 3 months has shown 0.83 precision, 0.74 recall rate, and detected 349 hidden vulnerabilities, proving the effectiveness and generality of the proposed approach.