Code Duplication and Reuse in Jupyter Notebooks

Code Duplication and Reuse in Jupyter Notebooks
复制标题

Jupyter Notebook 中的代码复制和重用

DOI:
--
复制
发表时间:
2020
期刊:
IEEE Symposium on Visual Languages / Human-Centric Computing Languages and Environments
影响因子:
--
通讯作者:
M. Storey
M. Storey
中科院分区:
--
文献类型:
--
作者:
Andreas Koenzen;Neil A. Ernst;M. Storey

文献摘要

参考文献

被引文献

相似文献

复制自己的代码可以加快编写软件的速度。这种便利对于计算机笔记本的用户来说特别有价值。复制允许笔记本用户快速测试假设并迭代数据。在本文中,我们探讨了计算笔记本中代码复制的数量、方式和来源,并确定了代码重用的潜在障碍。以前在计算笔记本领域的工作描述了开发人员重用和复制的动机,但没有显示重用发生了多少,或者他们在重用代码时面临哪些障碍。为了解决这个问题,我们首先分析了GitHub存储库中包含的Jupyter笔记本中的重复代码,然后对代码重用进行了观察性用户研究,参与者使用笔记本解决特定任务。我们的研究结果表明,我们样本中的存储库的平均自我复制率为7.6%。然而,在我们的用户研究中,很少有参与者复制他们自己的代码,他们更喜欢重用来自在线资源的代码。
Duplicating one’s own code makes it faster to write software. This expediency is particularly valuable for users of computational notebooks. Duplication allows notebook users to quickly test hypotheses and iterate over data. In this paper, we explore how much, how and from where code duplication occurs in computational notebooks, and identify potential barriers to code reuse. Previous work in the area of computational notebooks describes developers’ motivations for reuse and duplication but does not show how much reuse occurs or which barriers they face when reusing code. To address this gap, we first analyzed GitHub repositories for code duplicates contained in a repository’s Jupyter notebooks, and then conducted an observational user study of code reuse, where participants solved specific tasks using notebooks. Our findings reveal that repositories in our sample have a mean self-duplication rate of 7.6%. However, in our user study, few participants duplicated their own code, preferring to reuse code from online sources.
DOI: 10.1109/mc.2007.421
发表时间: 2007-12-01
期刊: COMPUTER
影响因子: 2.2
作者:
Gil, Yolanda;Deelman, Ewa;Myers, Jim
通讯作者: Myers, Jim