Casual Notebooks and Rigid Scripts: Understanding Data Science Programming

Casual Notebooks and Rigid Scripts: Understanding Data Science Programming
复制标题

休闲笔记本和僵化脚本:理解数据科学编程

DOI:
--
复制
发表时间:
2020
期刊:
IEEE Symposium on Visual Languages / Human-Centric Computing Languages and Environments
影响因子:
--
通讯作者:
Jan O. Borchers
Jan O. Borchers
中科院分区:
--
文献类型:
--
作者:
K. Subramanian;N. Hamdan;Jan O. Borchers

文献摘要

参考文献

被引文献

相似文献

数据工作者是非专业的数据科学家,他们经常使用R、Python或MATLAB等脚本语言,并采用探索性编程工作流。当前的IDE为他们提供了两种主要的编程模式:脚本文件和计算笔记本。为了了解这些模式如何影响工作实践,我们对21名数据工作者进行了一项研究,随后对62名受访者进行了更大规模的调查。通过访谈、步行和屏幕记录,我们收集了有关他们工作流程的信息。我们的分析表明,脚本和计算笔记本之间的紧张关系。重复更常见,更好地支持存储和执行以前的分析,但会阻碍实验。笔记本更适合实际的数据科学工作流程,但很容易变得杂乱无章。我们讨论了这种双重性质的模态使用如何导致影响数据工作者的工作流程的几个问题,并讨论了编程IDE的设计的影响。
Data workers are non-professional data scientists who often use scripting languages like R, Python, or MATLAB, and employ an exploratory programming workflow. Current IDEs offer them two main programming modalities: script files and computational notebooks. To understand how these modalities impact work practice, we conducted a study with 21 data workers, and a subsequent larger survey with 62 respondents. Through interviews, walkthroughs, and screen recordings, we collected information about their workflows. Our analysis shows a tension between scripts and computational notebooks. Scripts are more common, better support storage and execution of previous analyses, but hinder experimentation. Notebooks better suit the actual data science workflow, but can become easily unorganized. We discuss how this dual nature of modality usage leads to several issues that affect data workers’ workflows, and discuss implications for the design of programming IDEs.
通过带注释的单元折叠帮助计算笔记本的协作重用
DOI: 10.1145/3274419
发表时间: 2018
影响因子: --
作者:
Rule, Adam;Drosos, Ian;Tabard, Aurélien;Hollan, James D.
通讯作者: Hollan, James D.