Characteristics, potentials, and limitations of open-source Simulink projects for empirical research

Characteristics, potentials, and limitations of open-source Simulink projects for empirical research
复制标题

DOI:
10.1007/s10270-021-00883-0
复制
发表时间:
2021-04
影响因子:
2
通讯作者:
Alexander Boll;F. Brokhausen;Tiago Amorim;Timo Kehrer;Andreas Vogelsang
Alexander Boll;F. Brokhausen;Tiago Amorim;Timo Kehrer;Andreas Vogelsang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Alexander Boll;F. Brokhausen;Tiago Amorim;Timo Kehrer;Andreas Vogelsang

文献摘要

被引文献

相似文献

Simulink是将基于模型的开发范式成功应用于工业实践的一个例子。许多公司创建和维护Simulink项目,用于对软件密集型嵌入式系统进行建模,旨在实现早期验证和自动代码生成。然而,Simulink项目并不像基于代码的项目那样容易获得,基于代码的项目受益于大型可公开访问的开源存储库,从而抑制了实证研究。在本文中,我们调查了一组1734免费提供的Simulink模型从194个项目,并分析其适用性的实证研究。我们分析这些项目时考虑到(1)它们的开发背景,(2)它们在项目规模和组织方面的复杂性,以及(3)它们随时间的演变。我们的研究结果表明,实证研究既有局限性,也有潜力。一方面,一些应用领域主导着开发环境,并且有大量的模型可以被认为是有限的实际相关性的玩具示例。这些通常源于学术背景,仅由几个Simulink模块组成,并且不再(或从未)处于积极的开发或维护中。另一方面,我们发现,一个子集的分析模型是相当大的规模和复杂性。有些模型由数千个模块组成,其中一些模块由分层组织的Simulink子系统高度模块化。类似地,有些模型公开了几年的活动维护跨度,这表明它们在整个项目的生命周期中被用作主要的开发工件。根据与领域专家对我们的结果的讨论,许多模型可以被认为足够成熟以用于质量分析目的,并且它们暴露了可以被认为是行业规模模型的代表的特征。因此,我们相信,一个子集的模型是适合实证研究。更一般地说,使用公开可用的模型语料库或专用子集使研究人员能够复制研究结果,发表后续研究,并将其用于验证目的。我们发布我们的数据集是为了复制我们的结果并促进未来的实证研究。
Simulink is an example of a successful application of the paradigm of model-based development into industrial practice. Numerous companies create and maintain Simulink projects for modeling software-intensive embedded systems, aiming at early validation and automated code generation. However, Simulink projects are not as easily available as code-based ones, which profit from large publicly accessible open-source repositories, thus curbing empirical research. In this paper, we investigate a set of 1734 freely available Simulink models from 194 projects and analyze their suitability for empirical research. We analyze the projects considering (1) their development context, (2) their complexity in terms of size and organization within projects, and (3) their evolution over time. Our results show that there are both limitations and potentials for empirical research. On the one hand, some application domains dominate the development context, and there is a large number of models that can be considered toy examples of limited practical relevance. These often stem from an academic context, consist of only a few Simulink blocks, and are no longer (or have never been) under active development or maintenance. On the other hand, we found that a subset of the analyzed models is of considerable size and complexity. There are models comprising several thousands of blocks, some of them highly modularized by hierarchically organized Simulink subsystems. Likewise, some of the models expose an active maintenance span of several years, which indicates that they are used as primary development artifacts throughout a project’s lifecycle. According to a discussion of our results with a domain expert, many models can be considered mature enough for quality analysis purposes, and they expose characteristics that can be considered representative for industry-scale models. Thus, we are confident that a subset of the models is suitable for empirical research. More generally, using a publicly available model corpus or a dedicated subset enables researchers to replicate findings, publish subsequent studies, and use them for validation purposes. We publish our dataset for the sake of replicating our results and fostering future empirical research.