Variability Models in Large-Scale Systems: A Study and a Reverse-Engineering Technique

Variability Models in Large-Scale Systems: A Study and a Reverse-Engineering Technique
复制标题

大型系统中的可变性模型:研究和逆向工程技术

DOI:
--
复制
发表时间:
2015
期刊:
Software Engineering & Management
影响因子:
--
通讯作者:
Sarah Nadi
Sarah Nadi
中科院分区:
--
文献类型:
--
作者:
T. Berger;Sarah Nadi

文献摘要

被引文献

相似文献

高度可配置的系统可以轻松地拥有数千个配置选项,以及复杂的配置约束。可变性模型--选项和约束的高级表示--促进了大型的、高度可配置的系统的开发。由于模型很难创建和维护,我们努力支持这两个活动,尽可能地使它们自动化。为此,我们提出了一个实证研究的真实世界的可变性模型,静态代码分析技术,支持逆向工程和一致性检查等模型。导论.定制软件系统变得越来越重要。一种常见的方法是引入配置选项,在构建或启动时设置,以根据特定需求定制系统(例如,硬件、功能、性能)。大型系统可以很容易地拥有数千个这样的选项,以及它们之间复杂的配置约束。在更高级别的表示中声明选项和约束,即所谓的可变性模型,可以解决复杂性问题。这些模型便于大型、高度可配置系统的开发、配置、验证和确认[BNR 14]。然而,由于约束的复杂性,创建和维护可变性模型是费力且容易出错的。为了改善这种情况,我们需要(i)增加对大规模系统中此类模型的经验理解,以及(ii)构思采用和维护模型的方法,最好使用自动化技术。这次演讲介绍了我们为实现这两个目标所做的工作。首先,我们研究真实世界的模型,它们的底层语言,并开发工具,使这些模型可用于现成的推理机[BSL 13]。其次,我们开发了静态分析技术,从源代码中提取约束(可变性模型的主要成分),并评估自动逆向工程完整可变性模型的可行性[NBKC 14,NBKC 15]。真实世界的可变性模型。构想模型逆向工程方法需要真实模型的语料库。不幸的是,由于保密问题,研究人员几乎无法获得商业模型。现有的模型库,如S. P. L. O.T.,只包含小的学术模型。因此,我们研究了12个行业实力,开源项目的系统软件领域,其中包括Linux内核和eCos嵌入式操作系统(功能模型)的可变性模型。我们分析和逆向工程的建模语言的形式语义,并考虑到他们的表现力,开发命题抽象,使这些模型提供给现成的推理机(SAT求解器)。我们的语言分析进一步揭示了在实践中发现的复杂的语义和建模概念在学术建模语言中没有涉及,在研究中很少考虑。我们的模型分析确定了一组大型和复杂的模型,理想的候选人,以评估可扩展的模型逆向工程和分析技术。挖掘配置约束。我们假设模型中声明的许多配置约束来自低级别的源代码限制。为了研究这一点,我们开发了可扩展的静态分析技术,从高度可配置的系统的代码库中提取约束。我们用我们的分析来研究配置约束的起源,并调查在何种程度上,我们可以从代码逆向工程和consistencycheck模型。我们的静态分析依赖于两个规则来导出约束。第一条规则表示每个配置都应该正确地构建;也就是说,它应该进行预处理、解析、类型检查和链接。第二条规则表示每个有效的配置应该产生一个词汇上唯一的系统;也就是说,没有两个配置会导致完全相同的配置系统(即,源代码)。由于一个天真的方法,分析所有可能的组合的选项,将无法扩展,我们扩展现有的variabilityaware基础设施,使我们能够分析系统的所有变种一次。我们的基础设施FaRCE(约束提取)是公开可用的[NBKC 14]。我们应用我们的分析,我们以前调查的高度可配置的系统与可变性模型(Linux内核,eCos,uClibc,Busybox),以评估我们的方法的准确性,并确定哪些约束是可恢复的代码。我们发现,我们的方法是高度准确的(我们的两个规则分别为93%和77%),我们可以恢复28%的现有约束[NBKC 15]。然后,提取的约束条件可用于使用我们先前开发的算法合成实际变异性模型[SLB 11]。最后,我们定性地调查未恢复的约束的样本,确定15%的约束需要进一步分析(例如,数据流、控制流或动态分析),16%由于我们的工具的限制而不能被恢复,并且至少20%的约束是纯领域知识,这表明创建完整的模型需要进一步的实质性领域知识和测试。
Highly configurable systems can easily have thousands of configuration options, together with intricate configuration constraints. Variability models—higherlevel representations of options and constraints—facilitate the development of large, highly configurable systems. Since models are difficult to create and to maintain, we strive to support both activities, automating them as much as possible. To this end, we present an empirical study of real-world variability models, and static code-analysis techniques that support reverse-engineering and consistency-checking of such models. Introduction. Customizing software systems is becoming increasingly important. A common approach is to introduce configurations options, set at build or startup time, to tailor a system to specific needs (e.g., hardware, functionality, performance). Large systems can easily have thousands of such options, together with intricate configuration constraints among them. Declaring options and constraints in higher-level representations, so-called variability models, tackles complexity. Such models facility development, configuration, and validation and verification of large, highly configurable systems [BNR14]. However, creating and maintaining variability models is laborious and error-prone, given the complexity of constraints. To improve the situation, we need to (i) increase the empirical understanding of such models in large-scale systems, and (ii) conceive approaches to adopt and maintain models, ideally using automated techniques. This talk presents our work addressing these two goals. First, we study real-world models, their underlying languages, and develop tools to make these models available to off-the-shelf reasoners [BSL13]. Second, we develop static analysis techniques to extract constraints (the main ingredients of variability models) from source code and evaluate the feasibility of automatically reverse-engineering complete variability models [NBKC14, NBKC15]. Real-World Variability Models. Conceiving a model reverse-engineering approach requires a corpus of realistic models. Unfortunately, commercial models are hardly available to researchers, due to confidentiality issues. Existing model repositories, such as S.P.L.O.T., contain only small academic models. Thus, we study the (feature-model-like) variability models of twelve industry-strength, open-source projects from the systems software domain, including among others the Linux kernel and the eCos embedded operating system. We analyze and reverse-engineer the formal semantics of their modeling languages and, given their expressiveness, develop propositional abstractions to make these models available to off-the-shelf reasoners (SAT solvers). Our language analysis further reveals that there are intricate semantics and modeling concepts found in practice that are not covered in academic modeling languages and rarely considered in research. Our model analysis identifies a set of large and complex models–ideal candidates to evaluate scalable model reverse-engineering and analysis techniques. Mining Configuration Constraints. We hypothesize that many of the configuration constraints declared in the models arise from low-level source-code restrictions. To investigate this, we develop scalable static analysis techniques to extract constraints from the codebase of highly configurable systems. We use our analysis to study the origin of configuration constraints and to investigate to what extent we can reverse-engineer and consistencycheck models from code. Our static analysis relies on two rules to derive constraints. The first rule expresses that every configuration should build correctly; that is, it should pre-process, parse, type-check, and link. The second rule expresses that each valid configuration should yield a lexically unique system; that is, no two configurations lead to exactly the same configured system (i.e., source code). Since a naive approach—analyzing all possible combinations of options—would not scale, we extend an existing variabilityaware infrastructure which allows us to analyze all variants of the system at once. Our infrastructure FaRCE (FeatuRe Constraints Extraction) is publicly available [NBKC14]. We apply our analysis on four of our previously investigated highly configurable systems with variability models (Linux kernel, eCos, uClibc, Busybox) to evaluate the accuracy of our approach and to determine which constraints are recoverable from the code. We find that our approach is highly accurate (93 % and 77 % respectively for our two rules) and that we can recover 28 % of existing constraints [NBKC15]. The extracted constraints can then be used to synthesize actual variability models using our previously developed algorithms [SLB11]. Finally, we qualitatively investigate samples of non-recovered constraints, determining that 15 % of the constraints require further analyses (e.g., data-flow, control-flow, or dynamic analyses), 16 % could not be recovered due to limits of our tooling, and that at least 20 % of the constraints are purely domain knowledge—indicating that creating a complete model requires further substantial domain knowledge and testing.