A large-scale longitudinal study of flaky tests

A large-scale longitudinal study of flaky tests
复制标题

DOI:
10.1145/3428270
复制
发表时间:
2020-11
影响因子:
--
通讯作者:
Wing Lam;Stefan Winter;Anjiang Wei;Tao Xie;D. Marinov;Jonathan Bell
Wing Lam;Stefan Winter;Anjiang Wei;Tao Xie;D. Marinov;Jonathan Bell
中科院分区:
--
文献类型:
--
作者:
Wing Lam;Stefan Winter;Anjiang Wei;Tao Xie;D. Marinov;Jonathan Bell

文献摘要

被引文献

相似文献

片状测试是在回归测试效率下进行非确定性通过或失败的测试,因为开发人员无法轻易识别由于其最近的更改或由于理想情况而导致的。当引入片状时,可以进行测试,以便开发人员可以立即删除一些软件组织,例如Mozilla和Netflix某些工具 - 探测器 - 尽快检测片状测试,但是检测片状测试是由于它们固有的非确定性而造成的,因此即使是最先进的探测器也经常不切实际地用于所有测试项目更改以应对应用检测器的高成本,这些组织通常仅根据新添加或直接修改的测试来运行检测器对测试套件的更改,正在测试的代码和库依赖性。尺度对片状测试的纵向研究,以确定片状测试何时变片,什么变化会导致它们变得片状。确定可以在添加每个测试的代码版本中汇编并运行的245个片状测试。但是,在新添加的测试中,探测器仅在新添加的测试上运行,但仍会错过25%的片状测试。当检测器在新添加或直接修改的测试上运行时,将其增加到85%。评估测试何时变得片状,并建议将来应用检测器的指南。
Flaky tests are tests that can non-deterministically pass or fail for the same code version. These tests undermine regression testing efficiency, because developers cannot easily identify whether a test fails due to their recent changes or due to flakiness. Ideally, one would detect flaky tests right when flakiness is introduced, so that developers can then immediately remove the flakiness. Some software organizations, e.g., Mozilla and Netflix, run some tools—detectors—to detect flaky tests as soon as possible. However, detecting flaky tests is costly due to their inherent non-determinism, so even state-of-the-art detectors are often impractical to be used on all tests for each project change. To combat the high cost of applying detectors, these organizations typically run a detector solely on newly added or directly modified tests, i.e., not on unmodified tests or when other changes occur (including changes to the test suite, the code under test, and library dependencies). However, it is unclear how many flaky tests can be detected or missed by applying detectors in only these limited circumstances. To better understand this problem, we conduct a large-scale longitudinal study of flaky tests to determine when flaky tests become flaky and what changes cause them to become flaky. We apply two state-of-theart detectors to 55 Java projects, identifying a total of 245 flaky tests that can be compiled and run in the code version where each test was added. We find that 75% of flaky tests (184 out of 245) are flaky when added, indicating substantial potential value for developers to run detectors specifically on newly added tests. However, running detectors solely on newly added tests would still miss detecting 25% of flaky tests. The percentage of flaky tests that can be detected does increase to 85% when detectors are run on newly added or directly modified tests. The remaining 15% of flaky tests become flaky due to other changes and can be detected only when detectors are always applied to all tests. Our study is the first to empirically evaluate when tests become flaky and to recommend guidelines for applying detectors in the future.