Discovering Repetitive Code Changes in Python ML Systems

Discovering Repetitive Code Changes in Python ML Systems
复制标题

DOI:
10.1145/3510003.3510225
复制
发表时间:
2022-05
期刊:
2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Malinda Dilhara;Ameya Ketkar;Nikhith Sannidhi;Danny Dig
Malinda Dilhara;Ameya Ketkar;Nikhith Sannidhi;Danny Dig
中科院分区:
其他
文献类型:
--
作者:
Malinda Dilhara;Ameya Ketkar;Nikhith Sannidhi;Danny Dig

文献摘要

被引文献

相似文献

多年来,研究人员利用软件变化的重复性来自动化许多软件演变任务。尽管基于Python的ML系统的普及率大大增加,但它们并没有从这些进步中受益。不知道ML开发人员所做的重复变化是什么,研究人员,工具和图书馆设计师错过了自动化的机会,而ML开发人员未能学习和使用最佳编码实践。为了填补知识差距并推进了ML软件演化中的科学和工具,我们对1000个最高评级的ML系统的代码更改模式进行了第一项也是最精细的研究,其中包括5800万个SLOC。为了进行这项研究,我们将重复使用,适应和改进最先进的重复变化挖掘技术。我们的新型工具R-Cpatminer在400万以上的矿山挖掘并构造了350K细粒变化图,并检测28K变化模式。使用主题分析,我们确定了22个模式组,并揭示了ML开发人员如何更改其代码的4个主要趋势。我们调查了650台ML开发人员,以进一步阐明这些模式及其应用,并获得了15%的回应率。我们提出了对四个受众的可行,经验上的含义:(i)研究人员,(ii)工具构建者,(iii)ML图书馆供应商以及(iv)开发人员和教育者。
Over the years, researchers capitalized on the repetitiveness of software changes to automate many software evolution tasks. Despite the extraordinary rise in popularity of Python-based ML systems, they do not benefit from these advances. Without knowing what are the repetitive changes that ML developers make, researchers, tool, and library designers miss opportunities for automation, and ML developers fail to learn and use best coding practices. To fill the knowledge gap and advance the science and tooling in ML software evolution, we conducted the first and most fine-grained study on code change patterns in a diverse corpus of 1000 top-rated ML systems comprising 58 million SLOC. To conduct this study we reuse, adapt, and improve upon the state-of-the-art repetitive change mining techniques. Our novel tool, R-CPATMINER, mines over 4M commits and constructs 350K fine-grained change graphs and detects 28K change patterns. Using thematic analysis, we identified 22 pattern groups and we reveal 4 major trends of how ML developers change their code. We surveyed 650 ML developers to further shed light on these patterns and their applications, and we received a 15% response rate. We present actionable, empirically-justified implications for four audiences: (i) researchers, (ii) tool builders, (iii) ML library vendors, and (iv) developers and educators.