What Does It Take to Develop a Million Lines of Open Source Code?

What Does It Take to Develop a Million Lines of Open Source Code?
复制标题

开发一百万行开源代码需要什么?

DOI:
10.1007/978-3-642-02032-2_16
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
T. Mens
T. Mens
中科院分区:
--
文献类型:
--
作者:
J. Fernández;Daniel Izquierdo;T. Mens

文献摘要

被引文献

相似文献

本文对11个自由/Libre/开放源码软件(FLOSS)项目的规模与工作量、持续时间和团队规模之间的关系进行了初步和探索性的研究,这些项目目前的规模在60万到530万行代码(MLOC)之间。工作是根据每月活跃提交者的数量进行的。提取的数据与用于专有软件的早期版本的封闭源码成本估算模型Cocomo不太吻合,总体上表明,至少在某种程度上,FLOSS社区比封闭源码团队更有生产力。这也促使了对牙线专用努力模型的需求。作为第一近似值,我们评估了涉及不同属性对的16个线性回归模型。我们的实验之一是计算网络大小,即去除增长趋势中任何可疑的大异常值或跳跃。我们发现的最好的模型涉及努力与净规模,占差异的79%。该模型是基于排除了可能的异常值(Eclipse)的数据的,该数据是我们样本中最大的项目。这表明某些类别的FLOSS项目可能需要不同的努力模型。顺便说一句,对于11个单独的FLOSS项目中的每个项目,我们都能够以非常高的精度对净规模趋势进行建模(R2≥ 0.98.在11个项目中,3个呈超线性增长,5个呈线性增长,3个呈亚线性增长,这表明在大多数情况下,累积的复杂性要么得到了很好的控制,要么不构成增长的制约因素。
This article presents a preliminary and exploratory study of the relationship between size, on the one hand, and effort, duration and team size, on the other, for 11 Free/Libre/Open Source Software (FLOSS) projects with current size ranging between between 0.6 and 5.3 million lines of code (MLOC). Effort was operationalised based on the number of active committers per month. The extracted data did not fit well an early version of the closed-source cost estimation model COCOMO for proprietary software, overall suggesting that, at least to some extent, FLOSS communities are more productive than closed-source teams. This also motivated the need for FLOSS-specific effort models. As a first approximation, we evaluated 16 linear regression models involving different pairs of attributes. One of our experiments was to calculate thenet size, that is, to remove any suspiciously large outliers or jumps in the growth trends. The best model we found involved effort against net size, accounting for 79 percent of the variance. This model was based on data excluding a possible outlier (Eclipse), the largest project in our sample. This suggests that different effort models may be needed for certain categories of FLOSS projects. Incidentally, for each of the 11 individual FLOSS projects we were able to model the net size trends with very high accuracy (R2≥ 0.98). Of the 11 projects, 3 have grown superlinearly, 5 linearly and 3 sublinearly, suggesting that in the majority of the cases accumulated complexity is either well controlled or don’t constitute a growth constraining factor.