Model-Parallel Collaborative Filtering in Apache Spark
Model-Parallel Collaborative Filtering in Apache Spark
批准号:
1555772
负责人:
Ameet Talwalkar
金额:
$6.88万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2016-08-31
中文摘要
随着数据规模和复杂性的快速增长,许多组织都渴望在使用分布式计算环境的海量数据集上训练协同过滤方法。例如,Netflix有数十万个在线节目可以推荐给它的数百万用户,Facebook有数百万用户,他们可能会在彼此之间形成新的联系。然而,领先的方法在分布式环境中引入了重大的算法挑战。PI建议研究一种新的算法,旨在有效地用于大规模数据科学应用。初步研究证明了这种方法的前景,PI建议正式描述算法的行为,进行广泛的经验评估,并将受此建议启发的想法纳入PI即将教授的在线课程中。协同过滤,特别是矩阵分解,是设计推荐系统的一种广泛使用的方法。然而,这些模型的大小随着用户和项目的数量线性增长,并且由于其高昂的通信成本,矩阵分解的主要方法在分布式环境中引入了重大挑战。PI建议研究一种为Apache Spark设计的新型模型并行算法,该算法利用底层数据的稀疏性来大幅减少这种通信负担。初步研究证明了这种方法的前景,PI建议正式描述算法的行为,进行广泛的经验评估,并在其他学习环境中更普遍地探索Spark中的模型并行范式。PI还将把受此提议启发的与模型并行相关的想法整合到即将在edX平台上教授的MOOC中。
英文摘要
With data rapidly growing in size and complexity, many organizations are eager to train collaborative filtering methods on massive datasets using distributed computing environments. For instance, Netflix has hundreds of thousands of online programs to recommend to its millions of users, and Facebook has millions of users who could potentially form new links between one another. However, leading methods introduce significant algorithmic challenges in the distributed setting. The PI proposes to study a novel algorithm designed to be efficient for large-scale data science applications. Preliminary studies demonstrate the promise of this method, and the PI proposes to formally characterize the algorithm's behavior, perform an extensive empirical evaluation, and incorporate ideas inspired by this proposal into an upcoming online course PI will be teaching.Collaborative filtering, and in particular matrix factorization, is a widely used method for devising recommender systems. However, the size of these models grows linearly with the number of users and items, and leading methods for matrix factorization introduce significant challenges in the distributed setting due to their high communication costs. The PI proposes to study a novel model-parallel algorithm designed for Apache Spark that leverages the sparsity of the underlying data to drastically reduce this communication burden. Preliminary studies demonstrate the promise of this method, and the PI proposes to formally characterize the algorithm's behavior, perform an extensive empirical evaluation, and explore the paradigm of model-parallelism in Spark more generally for other learning settings. The PI will also incorporate ideas related to model-parallelism inspired by this proposal into an upcoming MOOC that be taught on the edX platform.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Travel: NSF Student Travel Grant for the Sixth Conference on Machine Learning and Systems (MLSys 2023)
-
批准号:2325547
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2023
-
负责人:Ameet Talwalkar
-
依托单位:
CAREER: Foundations of Next-Generation Neural Architecture Search
-
批准号:2046613
-
项目类别:Continuing Grant
-
资助金额:$55.0万
-
财政年份:2021
-
负责人:Ameet Talwalkar
-
依托单位:
BIGDATA: F: Optimization in Federated Networks of Devices
-
批准号:1838017
-
项目类别:Standard Grant
-
资助金额:$99.94万
-
财政年份:2019
-
负责人:Ameet Talwalkar
-
依托单位:
SIFTER: A Systems Biology Platform for Protein Function Prediction
-
批准号:1122732
-
项目类别:Fellowship Award
-
资助金额:$24.0万
-
财政年份:2011
-
负责人:Ameet Talwalkar
-
依托单位:
国内基金
海外基金
强流低能加速器束流损失机理的Parallel PIC/MCC算法与实现
-
批准号:11805229
-
项目类别:青年科学基金项目
-
资助金额:27.0万元
-
批准年份:2018
-
负责人:张青鵾
-
依托单位: