FLEX: fixing flaky tests in machine learning projects by updating assertion bounds

FLEX: fixing flaky tests in machine learning projects by updating assertion bounds
复制标题

DOI:
10.1145/3468264.3468615
复制
发表时间:
2021-08
期刊:
Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering
影响因子:
--
通讯作者:
Saikat Dutta;A. Shi;Sasa Misailovic
Saikat Dutta;A. Shi;Sasa Misailovic
中科院分区:
其他
文献类型:
--
作者:
Saikat Dutta;A. Shi;Sasa Misailovic

文献摘要

相似文献

许多机器学习(ML)算法本质上是随机的 - 使用相同的输入的多个执行可能会产生略有不同的结果,从而影响开发人员如何编写检查这些ML算法的端到端质量,选择适当的阈值以比较获得的质量指标与参考结果是一项非直觉的任务,这可能会导致我们提出flex的情况。由于ML算法中的算法随机性而自动固定片状测试的第一个工具。可以最大程度地减少我们技术的实际输出质量和预期的输出质量。极值理论,根据几个运行观察到的输出值的尾巴分布,FLEX更新测试中使用的界限,或者根据所需的置信度选择了测试的数量级别。对开发人员进行测试。
Many machine learning (ML) algorithms are inherently random – multiple executions using the same inputs may produce slightly different results each time. Randomness impacts how developers write tests that check for end-to-end quality of their implementations of these ML algorithms. In particular, selecting the proper thresholds for comparing obtained quality metrics with the reference results is a non-intuitive task, which may lead to flaky test executions. We present FLEX, the first tool for automatically fixing flaky tests due to algorithmic randomness in ML algorithms. FLEX fixes tests that use approximate assertions to compare actual and expected values that represent the quality of the outputs of ML algorithms. We present a technique for systematically identifying the acceptable bound between the actual and expected output quality that also minimizes flakiness. Our technique is based on the Peak Over Threshold method from statistical Extreme Value Theory, which estimates the tail distribution of the output values observed from several runs. Based on the tail distribution, FLEX updates the bound used in the test, or selects the number of test re-runs, based on a desired confidence level. We evaluate FLEX on a corpus of 35 tests collected from the latest versions of 21 ML projects. Overall, FLEX identifies and proposes a fix for 28 tests. We sent 19 pull requests, each fixing one test, to the developers. So far, 9 have been accepted by the developers.