Identification of an influence network using ensemble-based filtering for Hawkes processes driven by count data

Identification of an influence network using ensemble-based filtering for Hawkes processes driven by count data
复制标题

DOI:
10.1016/j.physd.2023.133676
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
N. Santitissadeekorn;S. Delahaies;Lloyd D.J.B
N. Santitissadeekorn;S. Delahaies;Lloyd D.J.B
中科院分区:
其他
文献类型:
--
作者:
N. Santitissadeekorn;S. Delahaies;Lloyd D.J.B

文献摘要

相似文献

许多网络具有事件驱动的动态(例如通信、社交媒体和犯罪网络),其中在网络中的节点处发生的事件的平均速率根据网络中其他事件的发生而变化。具体地,与网络的节点相关联的事件可以增加其他节点处的事件的速率,这取决于它们的影响关系。因此,使用时间数据来揭示给定网络的方向性、时间依赖性、影响结构,同时即使在缺乏物理网络的知识时也量化不确定性,这是令人感兴趣的。通常,用于推断网络中影响力结构的方法需要物理网络的知识或只能推断小型网络结构。在本文中,我们模型的事件驱动的动态网络的多维霍克斯过程。然后,我们开发了一种新的基于合奏的过滤方法的时间序列的计数数据(即,数据,提供每单位时间的事件数为网络中的每个节点),不仅跟踪影响网络结构随着时间的推移,但也通过合奏传播近似的不确定性。该方法克服了现有方法中的几个缺陷,例如用于推断多维Hawkes过程的现有方法太慢,对于超过150个节点的任何网络都不实用,只能处理时间戳数据(即事件发生时的数据,而不是每个节点的事件数量),并且我们不需要物理网络开始。我们的方法是大规模并行化的,允许其用于推断大型网络(10000个节点)的影响结构。我们展示了我们的方法,使用合成和真实世界的电子邮件通信数据的大型网络。
Many networks have event-driven dynamics (such as communication, social media and criminal networks), where the mean rate of the events occurring at a node in the network changes according to the occurrence of other events in the network. In particular, events associated with a node of the network could increase the rate of events at other nodes, depending on their influence relationship. Thus, it is of interest to use temporal data to uncover the directional, time-dependent, influence structure of a given network while also quantifying uncertainty even when knowledge of a physical network is lacking. Typically, methods for inferring the influence structure in networks require knowledge of a physical network or are only able to infer small network structures. In this paper, we model event-driven dynamics on a network by a multidimensional Hawkes process. We then develop a novel ensemble-based filtering approach for a time-series of count data (ie, data that provides the number of events per unit time for each node in the network) that not only tracks the influence network structure over time but also approximates the uncertainty via ensemble spread. The method overcomes several deficiencies in existing methods such as existing methods for inferring multidimensional Hawkes processes are too slow to be practical for any network over∼ 50 nodes, can only deal with timestamp data (ie data on just when events occur not the number of events at each node), and that we do not need a physical network to start with. Our method is massively parallelizable, allowing for its use to infer the influence structure of large networks (∼ 10, 000 nodes). We demonstrate our method for large networks using both synthetic and real-world email communication data.