Testing, Validation, and Verification of Robotic and Autonomous Systems: A Systematic Review

Testing, Validation, and Verification of Robotic and Autonomous Systems: A Systematic Review
复制标题

DOI:
10.1145/3542945
复制
发表时间:
2022-06
影响因子:
4.4
通讯作者:
Hugo L. S. Araujo;M. Mousavi;M. Varshosaz
Hugo L. S. Araujo;M. Mousavi;M. Varshosaz
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hugo L. S. Araujo;M. Mousavi;M. Varshosaz

文献摘要

被引文献

相似文献

我们对机器人和自主系统(RAS)的测试,验证和验证进行系统文献综述。这篇评论的范围涵盖了经过同行评审的研究论文,提出,改进或评估测试技术,过程或工具,以解决RAS的系统级质量。我们的调查是根据以三个阶段结构的严格方法进行的。首先,我们利用一组26篇种子论文(由域专家选择)和SERP检验分类法来设计我们的搜索查询和(特定领域的)分类法。其次,我们在三个学术搜索引擎中进行了搜索,并将包含和排除标准应用于结果。我们分别利用相关工作和领域专家(50名学者和15位行业专家)来验证和完善搜索查询。结果,我们遇到了10,735项研究,其中包括195项,审查和编码。我们的目标是回答四个研究问题,与(1)模型类型,(2)系统性能和测试是否充分措施,(3)工具及其可用性,以及(4)适用性的证据,尤其是在工业环境中。我们分析了我们的编码结果,以识别域中的优势和差距,并向研究人员和从业人员提出建议。我们的发现表明,时间逻辑的变体最广泛地用于建模需求和属性,而状态机器和过渡系统的变体被广泛用于建模系统行为。其他常见模型涉及指定要求的认知逻辑和指定系统行为的信念揭示模型。除了时间和认识论之外,模型中捕获的其他方面涉及概率(例如,用于建模不确定性)和连续轨迹(例如,用于建模车辆动力学和运动学)。许多论文缺乏对其提议的技术,过程或工具的效率,有效性或适当性的严格度量。在提供效率,有效性或充分性的衡量标准的人中,大多数使用域 - 不合时宜的通用度量,例如故障数量,状态空间的大小或验证时间。通过发展特定于域的性能和充分性概念来解决这方面的研究差距的趋势。为每个领域定义广泛接受的严格绩效和充分性衡量标准是确定的研究差距。在工具方面,使用最广泛的工具是诸如Prism和Uppaal之类的模型检查器,以及诸如凉亭之类的仿真工具。 MATLAB/SIMULINK是该域中的另一个广泛使用的工具集。总体而言,在该领域发表的论文中,有非常有限的行业适用性证据。即使考虑到各种自主系统的合并基准,甚至还有一个差距。
We perform a systematic literature review on testing, validation, and verification of robotic and autonomous systems (RAS). The scope of this review covers peer-reviewed research papers proposing, improving, or evaluating testing techniques, processes, or tools that address the system-level qualities of RAS. Our survey is performed based on a rigorous methodology structured in three phases. First, we made use of a set of 26 seed papers (selected by domain experts) and the SERP-TEST taxonomy to design our search query and (domain-specific) taxonomy. Second, we conducted a search in three academic search engines and applied our inclusion and exclusion criteria to the results. Respectively, we made use of related work and domain specialists (50 academics and 15 industry experts) to validate and refine the search query. As a result, we encountered 10,735 studies, out of which 195 were included, reviewed, and coded. Our objective is to answer four research questions, pertaining to (1) the type of models, (2) measures for system performance and testing adequacy, (3) tools and their availability, and (4) evidence of applicability, particularly in industrial contexts. We analyse the results of our coding to identify strengths and gaps in the domain and present recommendations to researchers and practitioners. Our findings show that variants of temporal logics are most widely used for modelling requirements and properties, while variants of state-machines and transition systems are used widely for modelling system behaviour. Other common models concern epistemic logics for specifying requirements and belief-desire-intention models for specifying system behaviour. Apart from time and epistemics, other aspects captured in models concern probabilities (e.g., for modelling uncertainty) and continuous trajectories (e.g., for modelling vehicle dynamics and kinematics). Many papers lack any rigorous measure of efficiency, effectiveness, or adequacy for their proposed techniques, processes, or tools. Among those that provide a measure of efficiency, effectiveness, or adequacy, the majority use domain-agnostic generic measures such as number of failures, size of state-space, or verification time were most used. There is a trend in addressing the research gap in this respect by developing domain-specific notions of performance and adequacy. Defining widely accepted rigorous measures of performance and adequacy for each domain is an identified research gap. In terms of tools, the most widely used tools are well-established model-checkers such as Prism and Uppaal, as well as simulation tools such as Gazebo; Matlab/Simulink is another widely used toolset in this domain. Overall, there is very limited evidence of industrial applicability in the papers published in this domain. There is even a gap considering consolidated benchmarks for various types of autonomous systems.