Towards Practical Incremental Recomputation for Scientists: An Implementation for the Python Language

Towards Practical Incremental Recomputation for Scientists: An Implementation for the Python Language
复制标题

为科学家提供实用的增量重新计算:Python 语言的实现

DOI:
--
复制
发表时间:
2010
期刊:
Workshop on the Theory and Practice of Provenance
影响因子:
--
通讯作者:
D. Engler
D. Engler
中科院分区:
--
文献类型:
--
作者:
Philip J. Guo;D. Engler

文献摘要

被引文献

相似文献

计算科学家经常使用像Python这样的高级语言为数据分析脚本制作原型。为了加快执行时间,他们手动将脚本重构为多个阶段(独立的函数),并编写额外的代码将中间结果保存到磁盘,以避免在后续运行中重新计算。为了消除这种负担,我们对Python解释器进行了增强,使其能够自动将长时间运行的函数执行结果缓存(保存)到磁盘,管理代码编辑和保存结果之间的依赖关系,并在确保安全的情况下重用缓存结果而不是重新执行那些函数。在初次运行时,运行时间会减慢约20%,但后续运行可能会加快几个数量级。使用我们增强的解释器,科学家可以编写简单且易于维护的代码,这些代码在经过小幅编辑后也能快速运行,而无需学习任何新的编程语言或结构。
Computational scientists often prototype data analysis scripts using high-level languages like Python. To speed up execution times, they manually refactor their scripts into stages (separate functions) and write extra code to save intermediate results to disk in order to avoid recomputing them in subsequent runs. To eliminate this burden, we enhanced the Python interpreter to automatically memoize (save) the results of long-running function executions to disk, manage dependencies between code edits and saved results, and re-use memoized results rather than re-executing those functions when guaranteed safe to do so. There is a ∼20% run-time slowdown during the initial run, but subsequent runs can speed up by several orders of magnitude. Using our enhanced interpreter, scientists can write simple and maintainable code that also runs fast after minor edits, without having to learn any new programming languages or constructs.