Comparing test quality measures for assessing student-written tests

Comparing test quality measures for assessing student-written tests
复制标题

比较评估学生书面测试的测试质量措施

DOI:
10.1145/2591062.2591164
复制
发表时间:
2014
期刊:
Companion Proceedings of the 36th International Conference on Software Engineering
影响因子:
--
通讯作者:
Z. Shams
Z. Shams
中科院分区:
--
文献类型:
--
作者:
S. Edwards;Z. Shams

文献摘要

被引文献

相似文献

现在,许多教育工作者都在编程任务中包括软件测试活动,因此可以手工分级测试的适当方法来评估学生写的软件测试的质量。目前使用的最常见的测量是覆盖范围 - 在相应的软件测试中,对学生的代码中的数量(在语句,分支或某些组合方面)进行了限制。高估了一些研究人员的真实质量。对于测试的质量,有点洞察力。通过对观察到的虫子的预测,当学生编写的测试的能力与自然发生的学生产生的缺陷相反时,测量测试质量的方法表明,所有对每个学生的测试都对每个学生进行了测试。其他学生的解决方案是揭示了测试套件的基础错误的最有效的预测指标,在漏洞揭示能力和代码覆盖率或突变分析分数之间没有强大的相关性。
Many educators now include software testing activities in programming assignments, so there is a growing demand for appropriate methods of assessing the quality of student-written software tests. While tests can be hand-graded, some educators also use objective performance metrics to assess software tests. The most common measures used at present are code coverage measures—tracking how much of the student’s code (in terms of statements, branches, or some combination) is exercised by the corresponding software tests. Code coverage has limitations, however, and sometimes it overestimates the true quality of the tests. Some researchers have suggested that mutation analysis may provide a better indication of test quality, while some educators have experimented with simply running every student’s test suite against every other student’s program—an “all-pairs” strategy that gives a bit more insight into the quality of the tests. However, it is still unknown which one of these measures is more accurate, in terms of most closely predicting the true bug revealing capability of a given test suite. This paper directly compares all three methods of measuring test quality in terms of how well they predict the observed bug revealing capabilities of student-written tests when run against a naturally occurring collection of student-produced defects. Experimental results show that all-pairs testing—running each student’s tests against every other student’s solution—is the most effective predictor of the underlying bug revealing capability of a test suite. Further, no strong correlation was found between bug revealing capability and either code coverage or mutation analysis scores.