An analysis of data sets used to train and validate cost prediction systems

An analysis of data sets used to train and validate cost prediction systems
复制标题

用于训练和验证成本预测系统的数据集分析

DOI:
--
复制
发表时间:
2005
期刊:
ACM SIGSOFT Softw. Eng. Notes
影响因子:
--
通讯作者:
M. Jørgensen
M. Jørgensen
中科院分区:
--
文献类型:
--
作者:
C. Mair;M. Shepperd;M. Jørgensen

文献摘要

被引文献

相似文献

目标-建立一个图片的性质和类型的数据集被用来开发和评估不同的软件项目工作量预测系统。我们认为这是重要的,因为有越来越多的出版工作,旨在评估不同的预测方法。方法-我们进行了详尽的搜索,从1980年起,从三个软件工程期刊的研究论文,使用项目数据集比较成本预测系统。结果-这确定了总共50篇论文,使用,一次或多次,共有71个独特的项目数据集。我们观察到,一些更知名和更容易获得的数据集被反复使用,使它们具有潜在的不成比例的影响力。这类数据集往往也是最古老的数据集之一,可能存在过时的问题。我们还注意到,只有大约60%的数据集是公共领域的。最后,从研究论文中提取相关信息一直是耗时的,由于不同风格的演示文稿和层次的上下文information.Conclusions -第一,社会需要考虑的质量和适当的数据集正在使用,并不是所有的数据集都是平等的。其次,我们需要评估结果的呈现方式,以促进荟萃分析,以及标准方案是否合适。
OBJECTIVE - to build up a picture of the nature and type of data sets being used to develop and evaluate different software project effort prediction systems. We believe this to be important since there is a growing body of published work that seeks to assess different prediction approaches.METHOD - we performed an exhaustive search from 1980 onwards from three software engineering journals for research papers that used project data sets to compare cost prediction systems.RESULTS - this identified a total of 50 papers that used, one or more times, a total of 71 unique project data sets. We observed that some of the better known and easily accessible data sets were used repeatedly making them potentially disproportionately influential. Such data sets also tend to be amongst the oldest with potential problems of obsolescence. We also note that only about 60% of all data sets are in the public domain. Finally, extracting relevant information from research papers has been time consuming due to different styles of presentation and levels of contextural information.CONCLUSIONS - first, the community needs to consider the quality and appropriateness of the data set being utilised; not all data sets are equal. Second, we need to assess the way results are presented in order to facilitate meta-analysis and whether a standard protocol would be appropriate.