Semantic image fuzzing of AI perception systems

Semantic image fuzzing of AI perception systems
复制标题

AI感知系统的语义图像模糊

DOI:
10.1145/3510003.3510212
复制
发表时间:
2022
期刊:
ICSE '22: Proceedings of the 44th International Conference on Software Engineering
影响因子:
--
通讯作者:
Sullivan, Kevin
Sullivan, Kevin
中科院分区:
--
文献类型:
--
作者:
Woodlief, Trey;Elbaum, Sebastian;Sullivan, Kevin

文献摘要

相似文献

感知系统能够解释物理世界的原始传感器读数的自主系统。感知系统的测试旨在揭示可能导致系统故障的误解。然而,目前的测试方法是不够的。人工解释和注释真实世界输入数据的成本很高,因此手动测试套件往往很小。模拟现实的差距降低了基于模拟世界的测试结果的有效性。综合测试输入的方法没有提供相应的预期解释。为了解决这些限制,我们开发了semSensFuzz,这是一种基于测试用例语义突变的感知系统模糊测试新方法,该测试用例将真实世界的传感器读数与其地面实况解释配对。我们实施了我们的方法,以评估其可行性和潜力,以改善感知系统的软件测试。我们用它为五个最先进的感知系统生成了150,000个语义变异的图像输入。我们发现,它综合了新颖的和主观现实的图像输入测试,它发现的输入,显示指定和计算的解释之间的显着不一致。我们还发现,它产生这样的测试用例的成本是非常低的人工语义注释的真实世界的图像相比。
Perception systemsenable autonomous systems to interpret raw sensor readings of the physical world. Testing of perception systems aims to reveal misinterpretations that could cause system failures. Current testing methods, however, are inadequate. The cost of human interpretation and annotation of real-world input data is high, so manual test suites tend to be small. The simulation-reality gap reduces the validity of test results based on simulated worlds. And methods for synthesizing test inputs do not provide corresponding expected interpretations. To address these limitations, we developedsemSensFuzz, a new approach to fuzz testing of perception systems based on semantic mutation of test cases that pair real-world sensor readings with their ground-truth interpretations. We implemented our approach to assess its feasibility and potential to improve software testing for perception systems. We used it to generate 150,000 semantically mutated image inputs for five state-of-the-art perception systems. We found that it synthesized tests with novel and subjectively realistic image inputs, and that it discovered inputs that revealed significant inconsistencies between the specified and computed interpretations. We also found that it produced such test cases at a cost that was very low compared to that of manual semantic annotation of real-world images.