Case consistency: a necessary data quality property for software engineering data sets

Case consistency: a necessary data quality property for software engineering data sets
复制标题

DOI:
10.1145/2745802.2745820
复制
发表时间:
2015-04
期刊:
Proceedings of the 19th International Conference on Evaluation and Assessment in Software Engineering
影响因子:
--
通讯作者:
Passakorn Phannachitta;Akito Monden;J. Keung;Ken-ichi Matsumoto
Passakorn Phannachitta;Akito Monden;J. Keung;Ken-ichi Matsumoto
中科院分区:
其他
文献类型:
--
作者:
Passakorn Phannachitta;Akito Monden;J. Keung;Ken-ichi Matsumoto

文献摘要

相似文献

数据质量是任何实证研究的一个重要方面,因为从实证数据中得出的模型和/或分析结果的有效性本质上受到其质量的影响。在这项实证研究中,我们专注于数据一致性作为一个关键因素,影响预测模型的准确性,在软件工程。我们提出了一个软件度量称为案例不一致性水平(CIL)分析软件工程数据集内的冲突,利用概率统计项目的情况下,并计算冲突对的数量。结果表明,CIL能够作为一个度量来识别一致的数据集或不一致的数据集,这是有价值的建立鲁棒的预测模型。除了测量一致性的水平,CIL被证明是适用于预测是否从数据集建立的努力模型可以达到更高的精度,在软件工程的经验实验的一个重要指标。
Data quality is an essential aspect in any empirical study, because the validity of models and/or analysis results derived from an empirical data is inherently influenced by its quality. In this empirical study, we focus on data consistency as a critical factor influencing the accuracy of prediction models in software engineering. We propose a software metric called Cases Inconsistency Level (CIL) for analyzing conflicts within software engineering data sets by leveraging probability statistics on project cases and counting the number of conflicting pairs. The result demonstrated that CIL is able to be used as a metric to identify either consistent data sets or inconsistent data sets, which are valuable for building robust prediction models. In addition to measuring the level of consistency, CIL is proved to be applicable to predict whether or not an effort model built from data set can achieve higher accuracy, an important indicator for empirical experiments in software engineering.