How to Better Distinguish Security Bug Reports (Using Dual Hyperparameter Optimization)

How to Better Distinguish Security Bug Reports (Using Dual Hyperparameter Optimization)
复制标题

DOI:
10.1007/s10664-020-09906-8
复制
发表时间:
2019-11
影响因子:
4.1
通讯作者:
Rui Shu;Tianpei Xia;Jianfeng Chen;L. Williams;T. Menzies
Rui Shu;Tianpei Xia;Jianfeng Chen;L. Williams;T. Menzies
中科院分区:
计算机科学2区
文献类型:
--
作者:
Rui Shu;Tianpei Xia;Jianfeng Chen;L. Williams;T. Menzies

文献摘要

被引文献

相似文献

为了使公众不容易受到黑客的攻击,安全漏洞报告需要在被广泛讨论之前由一小群工程师处理。但是,学习如何区分安全错误报告和其他错误报告是具有挑战性的,因为它们可能很少出现。能够找到这种稀缺目标的数据挖掘方法需要大量的优化工作。本研究的目标是帮助实践者优化那些试图区分罕见的安全错误报告和其他错误报告的方法。我们提出的方法SWIFT是一种双重优化器,它同时优化了学习器和预处理器选项。由于这是一个很大的选项空间,SWIFT使用了一种叫做𝜖-dominancethat的技术来学习如何避免不会显著提高性能的操作。结果:当与最近最先进的结果(发表在TSE ' 18上的FARSEC)进行比较时,我们发现SWIFT对预处理器和学习器的双重优化比单独优化它们更有用。例如,在对Chromium数据集的安全漏洞报告的研究中,FARSEC和SWIFT的召回率中位数分别为15.7%和77.4%。另一个例子是,在Ambari项目的数据实验中,中位数召回率从21.5%提高到85.7% (FARSEC到SWIFT)。总的来说,我们的方法可以快速优化模型,实现比现有技术更好的召回。回忆率的增加与假阳性率的适度增加相关(中位数从8%增加到24%)。对于未来的工作,这些结果表明双重优化是既实用又有用的。
BackgroundIn order that the general public is not vulnerable to hackers, security bug reports need to be handled by small groups of engineers before being widely discussed. But learning how to distinguish the security bug reports from other bug reports is challenging since they may occur rarely. Data mining methods that can find such scarce targets require extensive optimization effort.GoalThe goal of this research is to aid practitioners as they struggle to optimize methods that try to distinguish between rare security bug reports and other bug reports.MethodOur proposed method, called SWIFT, is adual optimizerthat optimizesbothlearner and pre-processor options. Since this is a large space of options, SWIFT uses a technique called𝜖-dominancethat learns how to avoid operations that do not significantly improve performance.ResultWhen compared to recent state-of-the-art results (from FARSEC which is published in TSE’18), we find that the SWIFT’s dual optimization of both pre-processor and learner is more useful than optimizing each of them individually. For example, in a study of security bug reports from the Chromium dataset, the median recalls of FARSEC and SWIFT were 15.7% and 77.4%, respectively. For another example, in experiments with data from the Ambari project, the median recalls improved from 21.5% to 85.7% (FARSEC to SWIFT).ConclusionOverall, our approach can quickly optimize models that achieve better recalls than the prior state-of-the-art. These increases in recall are associated with moderate increases in false positive rates (from 8% to 24%, median). For future work, these results suggest that dual optimization is both practical and useful.