Evolving Readable String Test Inputs Using a Natural Language Model to Reduce Human Oracle Cost

Evolving Readable String Test Inputs Using a Natural Language Model to Reduce Human Oracle Cost
复制标题

DOI:
10.1109/icst.2013.11
复制
发表时间:
2013-03
期刊:
2013 IEEE Sixth International Conference on Software Testing, Verification and Validation
影响因子:
--
通讯作者:
S. Afshan;Phil McMinn;Mark Stevenson
S. Afshan;Phil McMinn;Mark Stevenson
中科院分区:
其他
文献类型:
--
作者:
S. Afshan;Phil McMinn;Mark Stevenson

文献摘要

被引文献

相似文献

自动化oracle的频繁不可用意味着,在实践中,检查软件行为经常是一项艰苦的手工任务。尽管人类神谕师参与的成本很高,但很少有研究调查如何使这个角色更容易、更省时。人工oracle成本的一个来源是机器生成的测试输入固有的不可读性。特别是,自动生成的字符串输入往往是难以阅读的任意字符序列。这使得测试用例难以理解,并且检查起来非常耗时。在本文中,我们提出了一种将自然语言模型纳入基于搜索的输入数据生成过程的方法,目的是提高生成字符串的人类可读性。我们进一步介绍了在17个开源Java案例研究中使用该技术生成的测试输入的人类研究。在10个案例研究中,参与者在评估使用语言模型产生的输入时记录了明显更快的时间,其中60%的时间具有中等到较大的效应。此外,研究发现,其中3个案例的测试输入评估的准确性也得到了显著提高。
The frequent non-availability of an automated oracle means that, in practice, checking software behaviour is frequently a painstakingly manual task. Despite the high cost of human oracle involvement, there has been little research investigating how to make the role easier and less time-consuming. One source of human oracle cost is the inherent unreadability of machine-generated test inputs. In particular, automatically generated string inputs tend to be arbitrary sequences of characters that are awkward to read. This makes test cases hard to comprehend and time-consuming to check. In this paper we present an approach in which a natural language model is incorporated into a search-based input data generation process with the aim of improving the human readability of generated strings. We further present a human study of test inputs generated using the technique on 17 open source Java case studies. For 10 of the case studies, the participants recorded significantly faster times when evaluating inputs produced using the language model, with medium to large effect sizes 60% of the time. In addition, the study found that accuracy of test input evaluation was also significantly improved for 3 of the case studies.