What's your choice?: learning the mixed multi-nomial

What's your choice?: learning the mixed multi-nomial
复制标题

你的选择是什么?:学习混合多项式

DOI:
10.1145/2591971.2592020
复制
发表时间:
2014
期刊:
Journal of the Australian Mathematical Society. Series A. Pure Mathematics and Statistics
影响因子:
--
通讯作者:
L. Voloch
L. Voloch
中科院分区:
--
文献类型:
--
作者:
A. Ammar;Sewoong Oh;Devavrat Shah;L. Voloch

文献摘要

被引文献

相似文献

使用从异源人群中收集的消费者数据计算选择的排名已成为任何现代消费者信息系统的必不可少的模块,例如Yelp,Netflix,Amazon和App商店(例如Google Play)。在这样的应用中,排名或建议算法需要以可扩展的方式准确地从嘈杂数据中提取有意义的信息。解决这一挑战的原则方法需要一个模型,该模型将观测值与建议决策和利用此模型的可进行推理算法联系起来。为此,我们将消费者产生的偏好数据抽象为嘈杂的,部分意识到其先天偏好,即对选择的顺序或排列。受塞缪尔森(参见揭示的偏好的公理)和麦克法登的开创性作品的启发(参见运输的离散选择模型),我们将人群的先天偏好型建模为所谓的多名logit(MMNL)模型的混合物。在此模型下,建议问题归结为(a)从人口数据中学习MMNL模型,(b)在混合物中找到AM MNL组件,这些组件紧密地代表了手头消费者的偏好,并且(c)推荐其他选择。根据其发现的组成部分,对她/他的排名很高。在这项工作中,我们解决了从部分偏好中学习MMNL模型的问题。我们确定任何算法的基本局限性,以学习这种模型,并提供条件,在这些条件下,简单,数据驱动的(非参数)算法可以有效地学习模型。所提出的算法与标量(或Star)评级的标准协作过滤具有令人愉悦的相似性,但在排列域中。这项工作推进了学习分布的最新范围,而不是排列(参见[2])以及在学习混合分布的背景下(参见[4])。
Computing a ranking over choices using consumer data gathered from a heterogenous population has become an indispensable module for any modern consumer information system, e.g. Yelp, Netflix, Amazon and app-stores like Google play. In such applications, a ranking or recommendation algorithm needs to extract meaningful information from noisy data accurately and in a scalable manner. A principled approach to resolve this challenge requires a model that connects observations to recommendation decisions and a tractable inference algorithm utilizing this model. To that end, we abstract the preference data generated by consumers as noisy, partial realizations of their innate preferences, i.e. orderings or permutations over choices. Inspired by the seminal works of Samuelson (cf. axiom of revealed preferences) and that of McFadden (cf. discrete choice models for transportation), we model the population's innate preferences as a mixture of the so called Multi-nomial Logit (MMNL) model. Under this model, the recommendation problem boils down to (a) learning the MMNL model from population data, (b) finding am MNL component within the mixture that closely represents the revealed preferences of the consumer at hand, and (c) recommending other choices to her/him that are ranked high according to thus found component. In this work, we address the problem of learning MMNL model from partial preferences. We identify fundamental limitations of any algorithm to learn such a model as well as provide conditions under which, a simple, data-driven (non-parametric) algorithm learns the model effectively. The proposed algorithm has a pleasant similarity to the standard collaborative filtering for scalar (or star) ratings, but in the domain of permutations. This work advances the state-of-art in the domain of learning distribution over permutations (cf. [2]) as well as in the context of learning mixture distributions (cf. [4]).