A multi-level semantic web for hard-to-specify domain concept, Pedestrian, in ML-based software

A multi-level semantic web for hard-to-specify domain concept, Pedestrian, in ML-based software
复制标题

DOI:
10.1007/s00766-021-00366-0
复制
发表时间:
2022-01
影响因子:
2.8
通讯作者:
H. Barzamini;Murtuza Shahzad;Hamed Alhoori;Mona Rahimi
H. Barzamini;Murtuza Shahzad;Hamed Alhoori;Mona Rahimi
中科院分区:
计算机科学2区
文献类型:
--
作者:
H. Barzamini;Murtuza Shahzad;Hamed Alhoori;Mona Rahimi

文献摘要

相似文献

机器学习(ML)算法广泛用于构建软件密集型系统,包括安全关键型系统。与传统的软件组件不同,机器学习组件(MLC)是使用ML算法构建的软件组件,通过概括他们在有限的一组收集的示例中发现的共同特征来学习他们的规范。虽然这种归纳性质克服了编程的局限性,具体的概念,同样的功能成为问题,验证安全性的ML为基础的软件系统。一个原因是,由于MLC数据驱动的性质,通常没有一组明确编写和预定义的规范,MLC可以根据这些规范进行验证。在这方面,我们建议部分指定难以指定的领域概念,MLCs倾向于对其进行分类,而不是完全依赖于它们从任意收集的数据集中进行归纳学习的能力。在本文中,我们提出了一个半自动化的方法来构建一个多层次的语义网,部分概述了难以指定的,但至关重要的,在汽车领域的领域概念“行人”。我们以两种方式评估所生成的语义网的适用性:首先,参考网络,我们增加了一个行人数据集的缺失功能,轮椅,以显示在增强的数据集上训练一个最先进的基于ML的对象检测器提高了检测行人的准确性;第二,我们评估了基于多个最先进的行人和人类数据集的所生成的语义网的覆盖范围。
Machine Learning (ML) algorithms are widely used in building software-intensive systems, including safety-critical ones. Unlike traditional software components, Machine-Learned Components (MLC)s, software components built using ML algorithms, learn their specifications through generalizing the common features that they find in a limited set of collected examples. While this inductive nature overcomes the limitations of programminghard-to-specifyconcepts, the same feature becomes problematic for verifying safety in ML-based software systems. One reason is that, due to MLCs data-driven nature, there is often no set of explicitly written and pre-defined specifications, against which the MLC can be verified. In this regard, we propose to partially specify hard-to-specify domain concepts, which MLCs tend to classify, instead of fully relying on their inductive learning ability from arbitrarily-collected datasets. In this paper, we propose a semi-automated approach to construct a multi-level semantic web to partially outline the hard-to-specify, yet crucial, domain concept “pedestrian” in automotive domain. We evaluate the applicability of the generated semantic web in two ways: first, with a reference to the web, we augment a pedestrian dataset for a missing feature,wheelchair, to show training a state-of-the-art ML-based object detector on the augmented dataset improves its accuracy in detecting pedestrians; second, we evaluate the coverage of the generated semantic web based on multiple state-of-the-art pedestrian and human datasets.