In-database Distributed Machine Learning: Demonstration using Teradata SQL Engine

In-database Distributed Machine Learning: Demonstration using Teradata SQL Engine
复制标题

DOI:
10.14778/3352063.3352083
复制
发表时间:
2019-08-01
影响因子:
2.5
通讯作者:
Srivastava, Mani
Srivastava, Mani
中科院分区:
计算机科学2区
文献类型:
--
作者:
Sandha, Sandeep Singh;Cabrera, Wellington;Srivastava, Mani

文献摘要

被引文献

相似文献

机器学习促成了许多有趣的应用,并在大数据系统中得到广泛应用。流行的方法——在Tensorflow、Pytorch和Keras等框架中训练机器学习模型——需要将数据从数据库引擎移动到分析引擎,这给数据科学家增加了过多的开销,并成为模型训练的性能瓶颈。在本次演示中,我们实际展示了一种在数据库引擎内部原生实现分布式机器学习的解决方案。在演示过程中,观众将在Jupyter Notebook中交互使用Python API,直接在Teradata SQL引擎内部对合成回归数据集训练多元线性回归模型,并对视觉和感官数据集训练神经网络模型。
Machine learning has enabled many interesting applications and is extensively being used in big data systems. The popular approach - training machine learning models in frameworks like Tensorflow, Pytorch and Keras - requires movement of data from database engines to analytical engines, which adds an excessive overhead on data scientists and becomes a performance bottleneck for model training. In this demonstration, we give a practical exhibition of a solution for the enablement of distributed machine learning natively inside database engines. During the demo, the audience will interactively use Python APIs in Jupyter Notebooks to train multiple linear regression models on synthetic regression datasets and neural network models on vision and sensory datasets directly inside Teradata SQL Engine.