An In-depth Analysis of Duplicated Linux Kernel Bug Reports

An In-depth Analysis of Duplicated Linux Kernel Bug Reports
复制标题

DOI:
10.14722/ndss.2022.24159
复制
发表时间:
2022
期刊:
Proceedings 2022 Network and Distributed System Security Symposium
影响因子:
--
通讯作者:
Dongliang Mu;Yuhang Wu;Yueqi Chen;Zhenpeng Lin;Chensheng Yu;Xinyu Xing;Gang Wang
Dongliang Mu;Yuhang Wu;Yueqi Chen;Zhenpeng Lin;Chensheng Yu;Xinyu Xing;Gang Wang
中科院分区:
其他
文献类型:
--
作者:
Dongliang Mu;Yuhang Wu;Yueqi Chen;Zhenpeng Lin;Chensheng Yu;Xinyu Xing;Gang Wang

文献摘要

被引文献

相似文献

- 在过去的三年中,持续的fuzzing项目Syzkaller和Syzbot在检测内核漏洞方面取得了巨大的成功,发现的内核bug比过去20年发现的还要多。然而,连续模糊化的一个副作用是它会生成过多的崩溃报告,其中许多是由同一个bug引起的“重复”报告。虽然Syzbot使用简单的启发式方法对报告进行分组(重复数据删除),但我们发现它通常不准确。在本文中,我们经验性地分析了重复的内核错误报告,以了解:(1)重复的普遍性;(2)重复引入的潜在成本;(3)重复问题背后的关键原因。我们收集了从2017年9月到2020年11月的所有已修复内核错误,包括324万个崩溃报告,这些崩溃报告由Syzbot归类为2,526个错误报告(由独特的错误标题标识)。我们发现bug报告确实存在重复:2,526个bug报告中有47.1%与一个或多个其他报告重复。通过分析这些报告的元数据,我们发现未检测到的重复在时间和开发人员工作方面带来了额外的成本。然后,我们组织了Linux内核专家来分析一个重复的bug样本(375个bug报告,120个bug),并确定了导致重复的6个关键因素。基于这些实证研究结果,我们提出并原型化了可操作的错误重复数据删除策略。在使用地面实况数据集确认其有效性后,我们进一步应用我们的方法并识别出先前未知的重复
—In the past three years, the continuous fuzzing projects Syzkaller and Syzbot have achieved great success in detecting kernel vulnerabilities, finding more kernel bugs than those found in the past 20 years. However, a side effect of continuous fuzzing is that it generates an excessive number of crash reports, many of which are “duplicated” reports caused by the same bug. While Syzbot uses a simple heuristic to group (deduplicate) reports, we find that it is often inaccurate. In this paper, we empirically analyze the duplicated kernel bug reports to understand: (1) the prevalence of duplication; (2) the potential costs introduced by duplication; and (3) the key causes behind the duplication problem. We collected all of the fixed kernel bugs from September 2017 to November 2020, including 3.24 million crash reports grouped by Syzbot under 2,526 bug reports (identified by unique bug titles). We found the bug reports indeed had duplication: 47.1% of the 2,526 bug reports are duplicated with one or more other reports. By analyzing the metadata of these reports, we found undetected duplication introduced extra costs in terms of time and developer efforts. Then we organized Linux kernel experts to analyze a sample of duplicated bugs (375 bug reports, unique 120 bugs) and identified 6 key contributing factors to the duplication. Based on these empirical findings, we proposed and prototyped actionable strategies for bug deduplication. After confirming their effectiveness using a ground-truth dataset, we further applied our methods and identified previously unknown duplication