MLlib: Machine Learning in Apache Spark

MLlib: Machine Learning in Apache Spark
复制标题

DOI:
--
复制
发表时间:
2015-05
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Xiangrui Meng;Joseph K. Bradley;Burak Yavuz;Evan R. Sparks;S. Venkataraman;Davies Liu;Jeremy Freeman-Jeremy-Free
Xiangrui Meng;Joseph K. Bradley;Burak Yavuz;Evan R. Sparks;S. Venkataraman;Davies Liu;Jeremy Freeman-Jeremy-Free
中科院分区:
其他
文献类型:
--
作者:
Xiangrui Meng;Joseph K. Bradley;Burak Yavuz;Evan R. Sparks;S. Venkataraman;Davies Liu;Jeremy Freeman-Jeremy-Free

文献摘要

被引文献

相似文献

ApacheSpark是一个流行的大规模数据处理开源平台,非常适合迭代机器学习任务。本文介绍了Spark的开源分布式机器学习库MLlib。MLlib为广泛的学习设置提供了高效的功能,并包括几个基本的统计、优化和线性代数原语。与Spark一起提供的MLlib支持多种语言,并提供一个高级API,该API利用Spark丰富的生态系统来简化端到端机器学习管道的开发。MLlib经历了快速的增长,因为它有超过140名贡献者的充满活力的开源社区,并包括大量的文档来支持进一步的增长,并让用户快速掌握。
Apache Spark is a popular open-source platform for large-scale data processing that is well-suited for iterative machine learning tasks. In this paper we present MLlib, Spark's open-source distributed machine learning library. MLlib provides efficient functionality for a wide range of learning settings and includes several underlying statistical, optimization, and linear algebra primitives. Shipped with Spark, MLlib supports several languages and provides a high-level API that leverages Spark's rich ecosystem to simplify the development of end-to-end machine learning pipelines. MLlib has experienced a rapid growth due to its vibrant open-source community of over 140 contributors, and includes extensive documentation to support further growth and to let users quickly get up to speed.