CADE: The Missing Benchmark in Evaluating Dataset Requirements of AI-enabled Software

CADE: The Missing Benchmark in Evaluating Dataset Requirements of AI-enabled Software
复制标题

DOI:
10.1109/re54965.2022.00013
复制
发表时间:
2022-08
期刊:
2022 IEEE 30th International Requirements Engineering Conference (RE)
影响因子:
--
通讯作者:
H. Barzamini;Mona Rahimi
H. Barzamini;Mona Rahimi
中科院分区:
其他
文献类型:
--
作者:
H. Barzamini;Mona Rahimi

文献摘要

相似文献

由于这个原因,人工神经元模型的归纳性质使数据集质量成为其适当功能的关键因素研究通常缺乏可以评估所提出的质量指标的参考点。为了解释,评估和增强AI-Software数据集而建立可靠的参考点。 ,并评估与基准相对于基准的数据集语义质量和完整性。利用一系列新型的自然语言和图像处理技术来构建针对域规范的语义基准。规格,与数据概念的数据集差异和数据概念的代表性差异不足。 Cade产生的主题平均相关约75%。
The inductive nature of artificial neural models makes dataset quality a key factor of their proper functionality. For this reason, multiple research studies proposed metrics to assess the quality of the models’ datasets, such as dataset correctness, completeness, and consistency. However, these studies commonly lack a point of reference against which the proposed quality metrics could be assessed. To this end, this paper proposes a generic process that extracts the necessary knowledge to build a reliable reference point for the purpose of explanation, assessment, and augmentation of the AI-software dataset. This process automatically builds a benchmark specific to the software operational domain, interprets the training and validation datasets of AI-enabled perception software systems, and evaluates the dataset semantic quality and completeness relative to the benchmark. We implemented this process within a framework called Concept Augmentation and Dataset Evaluation (CADE), which leverages a series of novel natural language and image processing techniques to construct a semantic benchmark with respect to the domain specifications. The application of CADE to three commonly-used autonomous driving datasets showed several common weaknesses present in the arbitrarily-collected datasets against the encoded domain specifications, demonstrating dataset divergence from the domain concepts and under-represented variances of the concepts in the data. The qualitative evaluation results showed an average of about 75% relevancy of CADE generated topics.