Redundancy-free analysis of multi-revision software artifacts

Redundancy-free analysis of multi-revision software artifacts
复制标题

多版本软件工件的无冗余分析

DOI:
--
复制
发表时间:
2018
影响因子:
4.1
通讯作者:
H. Gall
H. Gall
中科院分区:
计算机科学2区
文献类型:
--
作者:
Carol V. Alexandru;Sebastiano Panichella;Sebastian Proksch;H. Gall

文献摘要

被引文献

相似文献

研究人员经常分析软件项目的一些修订,以获取有关其演变的历史数据。例如,他们通过静态分析源代码并监视某些指标在多个修订中的演变。运行这些分析的时间和资源要求通常使得有必要通过仅选择重大修订或使用粗粒度采样策略来限制分析的修订数量,例如,这可以消除进化的重要细节。大多数现有的分析技术并非用于分析多革命工件的设计,它们会单独处理每个修订。但是,两个随后的修订之间的实际差异通常很小。因此,量身定制的用于分析多个修订版的工具应仅分析这些差异,从而防止重新计算和存储冗余数据,提高可伸缩性并能够研究大量的修订。在这项工作中,我们提出了独立于精益语言的软件分析仪(LISA),这是一个通用框架,用于代表和分析多段的软件文物。它使用无冗余的多革命表示来进行工件,并仅通过分析数千个修订的伪影碎片来避免重新构成。对我们方法的评估包括衡量每种技术的效果,对LISA资源需求的深入研究以及对以四种语言编写的4,000个软件项目的700万个计划修订进行了大规模分析。我们表明,与传统的顺序方法相比,可以通过多个数量级来减少多个纠正分析的时间和空间要求。
Researchers often analyze several revisions of a software project to obtain historical data about its evolution. For example, they statically analyze the source code and monitor the evolution of certain metrics over multiple revisions. The time and resource requirements for running these analyses often make it necessary to limit the number of analyzed revisions, e.g., by only selecting major revisions or by using a coarse-grained sampling strategy, which could remove significant details of the evolution. Most existing analysis techniques are not designed for the analysis of multi-revision artifacts and they treat each revision individually. However, the actual difference between two subsequent revisions is typically very small. Thus, tools tailored for the analysis of multiple revisions should only analyze these differences, thereby preventing re-computation and storage of redundant data, improving scalability and enabling the study of a larger number of revisions. In this work, we propose the Lean Language-Independent Software Analyzer (LISA), a generic framework for representing and analyzing multi-revisioned software artifacts. It employs a redundancy-free, multi-revision representation for artifacts and avoids re-computation by only analyzing changed artifact fragments across thousands of revisions. The evaluation of our approach consists of measuring the effect of each individual technique incorporated, an in-depth study of LISA resource requirements and a large-scale analysis over 7 million program revisions of 4,000 software projects written in four languages. We show that the time and space requirements for multi-revision analyses can be reduced by multiple orders of magnitude, when compared to traditional, sequential approaches.