Data Mining Static Code Attributes to Learn Defect Predictors

Data Mining Static Code Attributes to Learn Defect Predictors
复制标题

DOI:
10.1109/tse.2007.256941
复制
发表时间:
2007
影响因子:
7.4
通讯作者:
T. Menzies;Jeremy Greenwald;A. Frank
T. Menzies;Jeremy Greenwald;A. Frank
中科院分区:
计算机科学1区
文献类型:
--
作者:
T. Menzies;Jeremy Greenwald;A. Frank

文献摘要

被引文献

相似文献

使用静态代码属性学习缺陷预测变量的价值已广泛争议。先前的工作已经探索了“麦卡布斯与霍尔斯特德与代码计数的行”的优点,以生成缺陷预测变量。我们在这里表明,此类辩论是无关紧要的,因为如何使用属性来构建预测因子要比使用特定属性更重要。同样,与先前的悲观主义相反,我们表明这种缺陷预测因子明显有用,并且在此处研究的数据上,平均检测概率为71%,平均虚假警报率为25%。这些预测因素对于优先考虑尚未检查的代码的优先级探索将是有用的。
The value of using static code attributes to learn defect predictors has been widely debated. Prior work has explored issues like the merits of "McCabes versus Halstead versus lines of code counts" for generating defect predictors. We show here that such debates are irrelevant since how the attributes are used to build predictors is much more important than which particular attributes are used. Also, contrary to prior pessimism, we show that such defect predictors are demonstrably useful and, on the data studied here, yield predictors with a mean probability of detection of 71 percent and mean false alarms rates of 25 percent. These predictors would be useful for prioritizing a resource-bound exploration of code that has yet to be inspected