On the use of mutation analysis for evaluating student test suite quality
On the use of mutation analysis for evaluating student test suite quality
复制标题
关于使用突变分析来评估学生测试套件的质量
DOI:
10.1145/3533767.3534217
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Bell, Jonathan
中科院分区:
文献类型:
--
作者:
Perretta, James;DeOrio, Andrew;Guha, Arjun;Bell, Jonathan
A common practice in computer science courses is to evaluate student-written test suites against either a set of manually-seeded faults (handwritten by an instructor) or against all other student-written implementations (“all-pairs” grading). However, manually seeding faults is a time consuming and potentially error-prone process, and the all-pairs approach requires significant manual and computational effort to apply fairly and accurately. Mutation analysis, which automatically seeds potential faults in an implementation, is a possible alternative to these test suite evaluation approaches. Although there is evidence in the literature that mutants are a valid substitute for real faults in large open-source software projects, it is unclear whether mutants are representative of the kinds of faults that students make. If mutants are a valid substitute for faults found in student-written code, and if mutant detection is correlated with manually-seeded fault detection and faulty student implementation detection, then instructors can instead evaluate student test suites using mutants generated by open-source mutation analysis tools.Using a dataset of 2,711 student assignment submissions, we empirically evaluate whether mutation score is a good proxy for manually-seeded fault detection rate and faulty student implementation detection rate. Our results show a strong correlation between mutation score and manually-seeded fault detection rate and a moderately strong correlation between mutation score and faulty student implementation detection. We identify a handful of faults in student implementations that, to be coupled to a mutant, would require new or stronger mutation operators or applying mutation operators to an implementation with a different structure than the instructor-written implementation. We also find that this correlation is limited by the fact that faults are not distributed evenly throughout student code, a known drawback of all-pairs grading. Our results suggest that mutants produced by open-source mutation analysis tools are of equal or higher quality than manually-seeded faults and a reasonably good stand-in for real faults in student implementations. Our findings have implications for software testing researchers, educators, and tool builders alike.
登录
查看更多内容
DOI:
10.1145/1029994.1029995
发表时间:
2003-09
期刊:
ACM J. Educ. Resour. Comput.
影响因子:
--
作者:
S. Edwards
通讯作者:
S. Edwards
DOI:
10.1145/2493394.2493402
发表时间:
2013
期刊:
Proceedings of the ninth annual international ACM conference on International computing education research
影响因子:
--
作者:
Z. Shams;S. Edwards
通讯作者:
S. Edwards
DOI:
--
发表时间:
2016
期刊:
International Symposium on Software Testing and Analysis
影响因子:
--
作者:
Henry Coles;Thomas Laurent;Christopher Henard;Mike Papadakis;Anthony Ventresque
通讯作者:
Anthony Ventresque
DOI:
10.1016/s0020-7373(83)80061-3
发表时间:
1983
期刊:
Int. J. Man Mach. Stud.
影响因子:
--
作者:
M. Weiser;Joan Shertz
通讯作者:
Joan Shertz
DOI:
10.1145/2591062.2591164
发表时间:
2014
期刊:
Companion Proceedings of the 36th International Conference on Software Engineering
影响因子:
--
作者:
S. Edwards;Z. Shams
通讯作者:
Z. Shams