Learning Mixtures of Plackett-Luce Models from Structured Partial Orders

Learning Mixtures of Plackett-Luce Models from Structured Partial Orders
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
--
影响因子:
--
通讯作者:
Zhibing Zhao;Lirong Xia
Zhibing Zhao;Lirong Xia
中科院分区:
其他
文献类型:
--
作者:
Zhibing Zhao;Lirong Xia

文献摘要

被引文献

相似文献

混合排序模型已被广泛用于异质偏好。然而,学习混合模型是非常不容易的,特别是当数据集由偏序组成的时候。在这种情况下,模型的参数可能甚至无法识别。本文主要研究了三种流行的偏序结构:排序最高的$L_1$、$L_2$-Way和一个备选方案子集上的选择数据。我们证明了当数据集由排序最高的$L_1$和$L_2$-Way的组合(或高达$L_2$备选方案的选择数据)组成时,当$L_1+L_2\le 2k-1$时,$k$Plackett-Luce模型的混合是不可识别的(当没有$L_2$-Way订单时,$L_2$被设置为$1$)。我们还证明了在某些组合下,包括排名前$3$、排名前$2$加$2$-way,以及高达$4$备选方案的选择数据,两个Plackett-Luce模型的混合是可识别的。在理论结果的指导下,我们提出了有效的广义矩方法(GMM)算法来学习两个Plackett-Luce模型的混合,并证明了它们是一致的。我们的实验证明了算法的有效性。此外,我们表明,当完整的排名可用时,从不同的边缘事件(偏序)学习提供了统计效率和计算效率之间的折衷。
Mixtures of ranking models have been widely used for heterogeneous preferences. However, learning a mixture model is highly nontrivial, especially when the dataset consists of partial orders. In such cases, the parameter of the model may not be even identifiable. In this paper, we focus on three popular structures of partial orders: ranked top-$l_1$, $l_2$-way, and choice data over a subset of alternatives. We prove that when the dataset consists of combinations of ranked top-$l_1$ and $l_2$-way (or choice data over up to $l_2$ alternatives), mixture of $k$ Plackett-Luce models is not identifiable when $l_1+l_2\le 2k-1$ ($l_2$ is set to $1$ when there are no $l_2$-way orders). We also prove that under some combinations, including ranked top-$3$, ranked top-$2$ plus $2$-way, and choice data over up to $4$ alternatives, mixtures of two Plackett-Luce models are identifiable. Guided by our theoretical results, we propose efficient generalized method of moments (GMM) algorithms to learn mixtures of two Plackett-Luce models, which are proven consistent. Our experiments demonstrate the efficacy of our algorithms. Moreover, we show that when full rankings are available, learning from different marginal events (partial orders) provides tradeoffs between statistical efficiency and computational efficiency.