Social Network and Click-through Prediction with Factorization Machines

Social Network and Click-through Prediction with Factorization Machines
复制标题

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Steffen Rendle
Steffen Rendle
中科院分区:
其他
文献类型:
--
作者:
Steffen Rendle

文献摘要

被引文献

相似文献

KDDCup 2012的两项任务是预测微博用户的关注者(主题1)和预测广告的点击率(主题2)。乍一看,这两个任务看起来不同,但它们都有两个重要的挑战。首先,两个问题设置中的主要变量都具有较大的分类域。用标准的机器学习模型估计这类变量之间的相互作用是困难的,因式分解模型已经成为这类数据的流行模型。其次,由于存在大量的预测变量,需要基于特征工程的灵活模型来简化模型定义。在这项工作中,它显示了如何分解机(FM)可以作为一种通用的方法来解决这两个轨道。FMs结合了特征工程的灵活性和分解模型的优点。本文简要介绍了FMs,并详细介绍了这两项任务的学习目标以及如何生成特征。对于轨道1,使用贝叶斯推理与马尔可夫链蒙特卡罗(MCMC),而对于轨道2,随机梯度下降(SGD)和基于MCMC的解决方案相结合。
The two tasks of KDDCup 2012 are to predict the followers of a microblogger (track 1) and to predict the click-through rate of ads (track 2). On rst glance, both tasks look dierent however they share two important challenges. First, the main variables in both problem settings are of large categorical domain. Estimating variable interactions between this type of variables is dicult with standard machine learning models and factorization models have become popular for such data. Secondly, many additional predictor variables are available which requires exible models based on feature engineering to facilitate model denition. In this work, it is shown how Factorization Machines (FM) can be used as a generic approach to solve both tracks. FMs combine the exibility of feature engineering with the advantages of factorization models. This paper shortly introduces FMs and presents for both tasks in detail the learning objectives and how features can be generated. For track 1, Bayesian inference with Markov Chain Monte Carlo (MCMC) is used whereas for track 2, stochastic gradient descent (SGD) and MCMC based solutions are combined.