Balancing Effectiveness and Flakiness of Non-Deterministic Machine Learning Tests

Balancing Effectiveness and Flakiness of Non-Deterministic Machine Learning Tests
复制标题

DOI:
10.1109/icse48619.2023.00154
复制
发表时间:
2023-05
期刊:
2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Chun Xia;Saikat Dutta;D. Marinov
Chun Xia;Saikat Dutta;D. Marinov
中科院分区:
其他
文献类型:
--
作者:
Chun Xia;Saikat Dutta;D. Marinov

文献摘要

相似文献

测试机器学习(ML)项目具有挑战性,因为各种ML算法具有固有的不确定性,并且缺乏可靠的方法来计算参考结果。开发人员在编写测试时通常依赖直觉来检查ML算法是否产生准确的结果。然而,这种方法导致在选择断言界限以比较测试断言中的实际结果和预期结果时的保守选择。由于开发人员希望避免测试中的假阳性失败,因此他们通常将边界设置得过于宽松,从而可能导致遗漏关键错误。我们提出了FASER -第一个系统的方法来平衡之间的权衡故障检测的有效性和片状的非确定性测试计算最佳断言界限。FASER通过改变断言界限将这种权衡框定为这些竞争目标之间的优化问题。FASER利用1)统计方法来估计剥落率,以及2)突变测试来估计故障检测有效性。我们评估了FASER从22个流行的ML项目中收集的87个非确定性测试。FASER发现,87个研究测试中有23个具有保守的界限,并提出了更严格的断言界限,最大限度地提高了测试的故障检测效率,同时限制了片状。我们已经向开发人员发送了19个拉取请求,每个请求修复一个测试,其中14个拉取请求已经被接受。
Testing Machine Learning (ML) projects is challenging due to inherent non-determinism of various ML algorithms and the lack of reliable ways to compute reference results. Developers typically rely on their intuition when writing tests to check whether ML algorithms produce accurate results. However, this approach leads to conservative choices in selecting assertion bounds for comparing actual and expected results in test assertions. Because developers want to avoid false positive failures in tests, they often set the bounds to be too loose, potentially leading to missing critical bugs. We present FASER - the first systematic approach for balancing the trade-off between the fault-detection effectiveness and flakiness of non-deterministic tests by computing optimal assertion bounds. FASER frames this trade-off as an optimization problem between these competing objectives by varying the assertion bound. FASER leverages 1) statistical methods to estimate the flakiness rate, and 2) mutation testing to estimate the fault-detection effectiveness. We evaluate FASER on 87 non-deterministic tests collected from 22 popular ML projects. FASER finds that 23 out of 87 studied tests have conservative bounds and proposes tighter assertion bounds that maximizes the fault-detection effectiveness of the tests while limiting flakiness. We have sent 19 pull requests to developers, each fixing one test, out of which 14 pull requests have already been accepted.