Automated Testing of Software that Uses Machine Learning APIs

Automated Testing of Software that Uses Machine Learning APIs
复制标题

DOI:
10.1145/3510003.3510068
复制
发表时间:
2022-05
期刊:
2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Chengcheng Wan;Shicheng Liu;Sophie Xie;Yifan Liu;H. Hoffmann;M. Maire;Shan Lu
Chengcheng Wan;Shicheng Liu;Sophie Xie;Yifan Liu;H. Hoffmann;M. Maire;Shan Lu
中科院分区:
其他
文献类型:
--
作者:
Chengcheng Wan;Shicheng Liu;Sophie Xie;Yifan Liu;H. Hoffmann;M. Maire;Shan Lu

文献摘要

相似文献

越来越多的软件应用程序结合了机器学习(ML)解决方案,用于统计模仿人类行为的认知任务。为了测试此类软件,需要巨大的人为努力来设计与软件相关的图像/文本/音频输入,并判断软件是否像大多数人类一样处理这些输入。即使暴露不当行为,罪魁祸首是在认知ML API内部还是使用API​​的代码。本文介绍了使用认知ML API的软件的新测试工具。饲养员为每个ML API设计一个伪内函数,该功能以经验方式逆转相应的认知任务(例如,图像搜索引擎伪转换图像分类API),并将这些伪内置功能结合到符号的执行引擎中自动发电相关图像/文本/音频输入并判断输出正确性。一旦暴露了不当行为,守护者将尝试改变软件中使用ML API来减轻行为不当的方式。我们对各种开源应用程序的评估表明,守护者可以大大改善分支机构的覆盖范围,同时识别许多前未知的错误。
An increasing number of software applications incorporate machine learning (ML) solutions for cognitive tasks that statistically mimic human behaviors. To test such software, tremendous human effort is needed to design image/text/audio inputs that are relevant to the software, and to judge whether the software is processing these inputs as most human beings do. Even when misbehavior is exposed, it is often unclear whether the culprit is inside the cognitive ML API or the code using the API. This paper presents Keeper, a new testing tool for software that uses cognitive ML APIs. Keeper designs a pseudo-inverse function for each ML API that reverses the corresponding cognitive task in an empirical way (e.g., an image search engine pseudo-reverses the image-classification API), and incorporates these pseudo-inverse functions into a symbolic execution engine to automatically gener-ate relevant image/text/audio inputs and judge output correctness. Once misbehavior is exposed, Keeper attempts to change how ML APIs are used in software to alleviate the misbehavior. Our evalu-ation on a variety of open-source applications shows that Keeper greatly improves the branch coverage, while identifying many pre-viously unknown bugs.