Maximizing data retention from the ISBSG repository

Maximizing data retention from the ISBSG repository
复制标题

最大限度地保留 ISBSG 存储库中的数据

DOI:
10.14236/ewic/ease2008.3
复制
发表时间:
2008
期刊:
--
影响因子:
--
通讯作者:
Stephen G. MacDonell
Stephen G. MacDonell
中科院分区:
--
文献类型:
--
作者:
Kefu Deng;Stephen G. MacDonell

文献摘要

被引文献

相似文献

背景:1997年,国际软件基准标准组(ISBSG)开始收集有关软件项目的数据。从那以后,他们通过一系列规模的发行序列向研究人员和从业人员提供了存储库的副本。 问题:关于存储库中数据的质量和完整性的问题,导致一些研究人员在观察结果方面丢弃了大量数据,并在模型中折现了软件开发工作的某些变量的使用。在某些情况下,丢弃数据的细节几乎没有提及,最少的理由。 方法:我们根据先前使用IFPUG/NESMA函数点分析(FPA)大小的项目(FPA)进行了大小并记录在项目级别的项目级别上,以最大程度地描述了用于在项目级别建模软件开发工作的数据量的过程中使用的过程。存储库。 结果:通过对数据集和域信息改进的正式形式化,我们得出了最终的可用数据集,其中包含13个变量的2862(3024)观察。 结论:一种有条理的数据进行预处理可以帮助确保保留尽可能多的数据以进行建模。假设数据确实反映了一个或多个基本模型,则这种保留应增加开发强大模型的可能性。
BACKGROUND: In 1997 the International Software Benchmarking Standards Group (ISBSG) began to collect data on software projects. Since then they have provided copies of their repository to researchers and practitioners, through a sequence of releases of increasing size. PROBLEM: Questions over the quality and completeness of the data in the repository have led some researchers to discard substantial proportions of the data in terms of observations, and to discount the use of some variables in the modelling of, among other things, software development effort. In some cases the details of the discarding of data has received little mention and minimal justification. METHOD: We describe the process we used in attempting to maximise the amount of data retained for modelling software development effort at the project level, based on previously completed projects that had been sized using IFPUG/NESMA function point analysis (FPA) and recorded in the repository. RESULTS: Through justified formalisation of the data set and domain-informed refinement we arrive at a final usable data set comprising 2862 (of 3024) observations across thirteen variables. CONCLUSION: a methodical approach to the pre-processing of data can help to ensure that as much data is retained for modelling as possible. Assuming that the data does reflect one or more underlying models, such retention should increase the likelihood of robust models being developed.