A Dirichlet Mixture Model of Hawkes Processes for Event Sequence Clustering

A Dirichlet Mixture Model of Hawkes Processes for Event Sequence Clustering
复制标题

DOI:
--
复制
发表时间:
2017-01
期刊:
--
影响因子:
--
通讯作者:
Hongteng Xu;H. Zha
Hongteng Xu;H. Zha
中科院分区:
其他
文献类型:
--
作者:
Hongteng Xu;H. Zha

文献摘要

被引文献

相似文献

基于一类特殊但重要的点过程-- Hawkes过程的Dirichlet混合模型,提出了一种有效的事件序列聚类方法.在该模型中,属于一个集群的每个事件序列都是通过具有特定参数的相同Hawkes过程生成的,不同的集群对应不同的Hawkes过程。Hawkes过程的先验分布由Dirichlet分布控制。我们通过最大似然估计(MLE)学习模型,并提出了一个有效的变分贝叶斯推理算法。我们具体分析了EM型算法的上下文中的内外迭代,并讨论了几个内部迭代分配策略。我们的模型的可识别性,我们的学习方法的收敛性,其样本的复杂性进行了分析,在理论和实证的方式,这表明我们的方法优于其他竞争对手。所提出的方法自动学习聚类的数量,并且对模型误指定具有鲁棒性。在合成数据和真实数据上的实验表明,该方法可以学习到异步事件序列中隐藏的不同触发模式,并在聚类纯度和一致性方面取得了令人鼓舞的性能。
We propose an effective method to solve the event sequence clustering problems based on a novel Dirichlet mixture model of a special but significant type of point processes --- Hawkes process. In this model, each event sequence belonging to a cluster is generated via the same Hawkes process with specific parameters, and different clusters correspond to different Hawkes processes. The prior distribution of the Hawkes processes is controlled via a Dirichlet distribution. We learn the model via a maximum likelihood estimator (MLE) and propose an effective variational Bayesian inference algorithm. We specifically analyze the resulting EM-type algorithm in the context of inner-outer iterations and discuss several inner iteration allocation strategies. The identifiability of our model, the convergence of our learning method, and its sample complexity are analyzed in both theoretical and empirical ways, which demonstrate the superiority of our method to other competitors. The proposed method learns the number of clusters automatically and is robust to model misspecification. Experiments on both synthetic and real-world data show that our method can learn diverse triggering patterns hidden in asynchronous event sequences and achieve encouraging performance on clustering purity and consistency.