Versioning for End-to-End Machine Learning Pipelines
Versioning for End-to-End Machine Learning Pipelines
复制标题
端到端机器学习管道的版本控制
DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
T. V. Kasteren
中科院分区:
文献类型:
--
作者:
T. V. D. Weide;D. Papadopoulos;O. Smirnov;Michal Zielinski;T. V. Kasteren
End-to-end machine learning pipelines that run in shared environments are challenging to implement. Production pipelines typically consist of multiple interdependent processing stages. Between stages, the intermediate results are persisted to reduce redundant computation and to improve robustness. Those results might come in the form of datasets for data processing pipelines or in the form of model coefficients in case of model training pipelines. Reusing persisted results improves efficiency but at the same time creates complicated dependencies. Every time one of the processing stages is changed, either due to code change or due to parameters change, it becomes difficult to find which datasets can be reused and which should be recomputed. In this paper we build upon previous work to produce derivations of datasets to ensure that multiple versions of a pipeline can run in parallel while minimizing the amount of redundant computations. Our extensions include partial derivations to simplify navigation and reuse, explicit support for schema changes of pipelines, and a central registry of running pipelines to coordinate upgrading pipelines between teams.